Fetching from the wire…
Models2026-08-21 · source-backed
The progression: 82, then ~114, then ~138 with DFlash2 drafting and lookup-augmented drafting, and now ~133 tok/s on real chat prompts with 382 tok/s when the model reproduces its own context. Stack is fp8 KV cache, int8 lm_head and embed_tokens, fp16 recurrent state, int8 activations, W4A16-requantized DFlash2 block drafting, prefix caching, split-KV verify attention, KVarN for 262k context. r/LocalLLaMA The number they care about is 15 of 16 tokens accepted per verify step on document-quoting workloads, which is the honest caveat: the 382 figure is a best case on a specific shape of work.
Each link below shares sources, entities, or timing with this story.
A user moved from UD-Q3_K_XL at 140k context to UD-IQ3_XXS and cleared 200k on a 16GB eGPU over Thunderbolt 4, with KV cache at q5_1 and llama.cpp built with DGGML_CUDA_FA_ALL_QUANTS=ON (r/LocalLLaMA). Prompt processing fell from 700-800 tok/s to 400. A commenter on an RTX 508...
Two published vLLM configs for a single 24GB card at 250W with 150k context: batch mode measuring ~1,094 tok/s steady-state decode at 64 concurrent (942 end-to-end, rising to ~1,222/1,042 with all layers int8), and single-user mode at 114-122 tok/s single-stream via MTP specul...
A developer published v100-skinny with hand-written NVFP4 W4A16 CUDA kernels plus chain-MTP speculative serving: four V100s at 219.1 ± 5.9 tok/s decode against a 5090 running NInfer at 214.7 ± 9.2, both 5/5 correct on AIME 2026 problem 1 across five seeds. (r/LocalLLaMA) The m...
Alibaba released Qwen3.6-27B on April 22. Dense architecture. Open weights. 77.2% on SWE-bench Verified, within 3.7 points of Claude Opus 4.6. On SkillsBench, it scores 48.2% versus its own 397B MoE predecessor's 30.0%. That's a 77% improvement with 14.8x fewer parameters. Let...
Alibaba released Qwen3.5-0.8B, 2B, 4B, and 9B — all natively multimodal (text+image+video from same weights, no adapter), 262K context, Apache 2.0. The 9B beats last-gen Qwen3-30B across the board and outperforms GPT-5-Nano by 13 points on MMMU-Pro. Architecture uses Gated Del...
For about a year, "run your agent locally" meant accepting a model that couldn't reliably call a tool twice in a row. That excuse is gone. Meta Superintelligence Labs published Muse Glimmer today: a 29.6B dense causal transformer, 52 layers, 6,656 hidden dim, with a ~1.8B ViT-...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.