Fetching from the wire…
Models2026-08-12 · source-backed
Following an earlier claim of 1,000 tok/s aggregate on the same hardware (r/LocalLLaMA). 106 upvotes, 90 comments, a 0.85 ratio indicating heavy scrutiny, which is warranted: V100 is Volta with no native FP4 path, so everything rides on the emulation approach. Single-source, unverified by any harness. If it holds it materially changes the economics of used-V100 clusters, which is a big if.
Each link below shares sources, entities, or timing with this story.
A developer published v100-skinny with hand-written NVFP4 W4A16 CUDA kernels plus chain-MTP speculative serving: four V100s at 219.1 ± 5.9 tok/s decode against a 5090 running NInfer at 214.7 ± 9.2, both 5/5 correct on AIME 2026 problem 1 across five seeds. (r/LocalLLaMA) The m...
A 100-task sweep with mini-SWE-agent 2.4.6 on an RTX PRO 6000 WS, sglang, NVFP4 weights, full 262K context, three templates at two reasoning efforts. Stock went 91% at medium and 99% at xhigh. Fixed went 87% to 98%. Sharp sat flat at 94% for both. The cost side decides it: sto...
Alibaba published a fine-grained MoE with 2.4T total / 95B active, 512 experts, and a 92-layer hybrid full/linear attention backbone. vLLM shipped day-0 support verified on NVIDIA and AMD with ready 4-bit checkpoints (NVFP4 at 1.32 TiB for an 8xB300 node, MXFP4 at 1.45 TiB for...
An r/LocalLLaMA post hit 809 upvotes and 143 comments on Daniel Han's claim that the unreleased 27B will run on 17GB RAM/VRAM setups. Unsloth's X account confirms the number verbatim. Alibaba has published no benchmark table, license, or activated-parameter count for the 27B....
TAK builds an imatrix from a task-specific corpus, finds the smallest size before collapse, then promotes and demotes tensors within a byte budget. No pruning, no fine-tuning, no merging. Held-out reasoning: 82.81% against 83.59% for BF16 and 77.34% for byte-matched Unsloth UD...
An RTX 5080 owner ranked community quantizations by mean KLD and same-top-p agreement rather than a public benchmark. bartowski/Qwen3.8-27B-IQ4_XS won overall, huihui-ai's abliterated UD-IQ4_XS was the best uncensored option, and jpetrina's IQ4_XS-pure is the pick when you nee...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.