Fetching from the wire…
Infra2026-08-20 · source-backed
Two published vLLM configs for a single 24GB card at 250W with 150k context: batch mode measuring ~1,094 tok/s steady-state decode at 64 concurrent (942 end-to-end, rising to ~1,222/1,042 with all layers int8), and single-user mode at 114-122 tok/s single-stream via MTP speculation with four cheap drafts, calibrated int4 lm_head and split-KV verify attention. (GitHub) The standout number: 381 tok/s at 25k context when the model reproduces its own context, quoting a document or applying an edit, using DFlash2 with 15 drafted tokens per verify step. Speculation wins below roughly 8 concurrent users, plain batching wins above.
Each link below shares sources, entities, or timing with this story.
The progression: 82, then ~114, then ~138 with DFlash2 drafting and lookup-augmented drafting, and now ~133 tok/s on real chat prompts with 382 tok/s when the model reproduces its own context. Stack is fp8 KV cache, int8 lm_head and embed_tokens, fp16 recurrent state, int8 act...
The repo appeared August 24, opening the weights of a multimodal MoE that Qwen frames explicitly as an architecture preview, the same role Qwen3-Next played for Qwen3.5 (GitHub). The hybrid Gated DeltaNet plus Gated Attention design it previews already carried through the Qwen...
Alibaba released Qwen3.6-27B on April 22. Dense architecture. Open weights. 77.2% on SWE-bench Verified, within 3.7 points of Claude Opus 4.6. On SkillsBench, it scores 48.2% versus its own 397B MoE predecessor's 30.0%. That's a 77% improvement with 14.8x fewer parameters. Let...
AlexsJones/llmfit released v1.1.10 today, adding RamaLama runtime discovery to its MCP server, the Qwen3.8 model family and MiniMax M3 vision capability exposure (GitHub). It also merged 32 MLX benchmark results on an Apple M4 Pro, the project's first MLX entries, giving an ap...
A developer published v100-skinny with hand-written NVFP4 W4A16 CUDA kernels plus chain-MTP speculative serving: four V100s at 219.1 ± 5.9 tok/s decode against a 5090 running NInfer at 214.7 ± 9.2, both 5/5 correct on AIME 2026 problem 1 across five seeds. (r/LocalLLaMA) The m...
1. Set Up Cursor Automations (intermediate) — Event-driven agents from PagerDuty/GitHub/Slack triggers with isolated sandboxes. Cursor Blog 2. Apply Context Engineering to Cut Agent Costs 60-80% (advanced) — Hierarchical token budgets, dynamic tool filtering (max 15), automati...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.