Fetching from the wire…
Infra2026-08-20 · source-backed
A developer published v100-skinny with hand-written NVFP4 W4A16 CUDA kernels plus chain-MTP speculative serving: four V100s at 219.1 ± 5.9 tok/s decode against a 5090 running NInfer at 214.7 ± 9.2, both 5/5 correct on AIME 2026 problem 1 across five seeds. (r/LocalLLaMA) The mechanism is the honest part: the V100 round takes 26.9ms vs 19.9ms (35% slower) but commits 5.89 tokens per round vs 4.27 (38% more), because the setup runs Qwen3.8's own MTP at depth 7 while NInfer caps at 5. Speculative depth bought back a hardware generation.
Each link below shares sources, entities, or timing with this story.
Following an earlier claim of 1,000 tok/s aggregate on the same hardware (r/LocalLLaMA). 106 upvotes, 90 comments, a 0.85 ratio indicating heavy scrutiny, which is warranted: V100 is Volta with no native FP4 path, so everything rides on the emulation approach. Single-source, u...
3B active parameters, beats Qwen3.5-35B-A3B on AIME 2025 (92.4 vs 91.9), LiveCodeBench v6 (87.2 vs 74.6), and surpasses the larger Nemotron-3-Super-120B. Available on Ollama and HuggingFace under open license. Source
Alibaba released Qwen3.6-27B on April 22. Dense architecture. Open weights. 77.2% on SWE-bench Verified, within 3.7 points of Claude Opus 4.6. On SkillsBench, it scores 48.2% versus its own 397B MoE predecessor's 30.0%. That's a 77% improvement with 14.8x fewer parameters. Let...
An r/LocalLLaMA post reports a working deployment on 80x RTX 5090 connected over 25 gigabit Ethernet rather than NVLink or InfiniBand — roughly 2.5TB aggregate VRAM against a ~594GB MXFP4 weight file, surplus absorbed by activation and KV overhead. Expert-parallel MoE over com...
A user moved from UD-Q3_K_XL at 140k context to UD-IQ3_XXS and cleared 200k on a 16GB eGPU over Thunderbolt 4, with KV cache at q5_1 and llama.cpp built with DGGML_CUDA_FA_ALL_QUANTS=ON (r/LocalLLaMA). Prompt processing fell from 700-800 tok/s to 400. A commenter on an RTX 508...
The repo appeared August 24, opening the weights of a multimodal MoE that Qwen frames explicitly as an architecture preview, the same role Qwen3-Next played for Qwen3.5 (GitHub). The hybrid Gated DeltaNet plus Gated Attention design it previews already carried through the Qwen...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.