Fetching from the wire…
Top 5 · 2026-08-02 · source-backed
DeepSeek-V4-Flash-0731 landed July 31 under MIT with a DSpark speculative-decoding module attached. Terminal Bench 2.1: 82.7. Toolathlon-Verified: 70.3. DSBench-FullStack: 68.7. DeepSWE: 54.4. NL2Repo: 54.2. The model card claims it beats DeepSeek-V4-Pro (Preview) "despite its far smaller activated parameter count," which is an unusual thing for a lab to say about its own flagship.
API pricing is $0.14 in / $0.28 out per million tokens with a 98% cache discount. For reference, CostPerPrompt currently lists GPT-5.6 Sol at $5/$30 and Claude Opus 5 at $5/$25. Two orders of magnitude.
It took #2 on Product Hunt August 1 with 299 upvotes, and it's the single most-discussed item across HN and r/LocalLLaMA today. But the useful reporting is coming from people running it, not rating it.
Pull the chat template fix before you judge this model. llama.cpp PR #26398, opened August 1 by tarruda, corrects the V4 preview Jinja template to match official encoder behavior and adds a dedicated 0731-variant template with distinct prompts per reasoning level, handling max-effort reasoning, structured output, and a drop_thinking default of True. The r/LocalLLaMA poster who flagged it reported looping and garbage tool-calling the day before that stopped completely after the fix landed. The model was fine. The template wasn't. If you pulled GGUFs last week and concluded V4 Flash is bad at agentic work, you evaluated a bug.
Two llama.cpp issues are still open and worth knowing before you commit. #26399 reports GGML_OP_TOP_K falling back to CPU on HIP/ROCm above roughly 3-4K context, costing a 6.4x token-generation loss precisely where agent contexts live. #26423 reports quantized KV cache still producing garbage on master.
Real hardware numbers: an r/LocalLLaMA member got roughly 3.5 tok/s at IQ2_M on dual RTX 3060s plus 96GB system RAM, with the sharp detail that LM Studio refused to distribute weights onto the second GPU while Unsloth Studio did. That's a loader difference that presents as an OOM failure. Meanwhile antirez's ds4, a single-file C inference engine at 19,834 stars (+150 today), widened from Metal-only to Metal, CUDA and ROCm and now covers V4 PRO as well as Flash. Third-party reviews report 26 tok/s at 50W peak on a 128GB M3 Max under his own Q2 quantization.
The uncomfortable read for anyone selling software: the cost basis you budgeted for agent inference last quarter is now the premium option. If your product's margin depends on inference being expensive, that assumption has a shelf life measured in months.
Each link below shares sources, entities, or timing with this story.
The changelog lists Terminal Bench 2.1 at 82.7, NL2Repo 54.2, Cybergym 76.7, DeepSWE 54.4, Toolathlon verified 70.3, DSBench-FullStack 68.7 and DSBench-Hard 59.6: figures DeepSeek says far exceed V4-Pro-Preview. Native Responses API support and specific Codex adaptation. Only...
For about a year, "run your agent locally" meant accepting a model that couldn't reliably call a tool twice in a row. That excuse is gone. Meta Superintelligence Labs published Muse Glimmer today: a 29.6B dense causal transformer, 52 layers, 6,656 hidden dim, with a ~1.8B ViT-...
The leaderboard says first place. The methodology says you should check your own bill. Qwen3.8 Max now ranks first on Artificial Analysis' agentic index, scoring 86.1 on OSWorld-Verified ahead of GPT-5.6 Sol Max at 83.2 and Fable 5 at 85.0, priced at $2.00/M input and $6.00/M...
DeepSeek released V4 on April 24 and the numbers demand attention. V4-Pro is 1.6 trillion parameters total with 49 billion active, MIT-licensed, native 1M-token context. It scores 80.6% on SWE-bench Verified, putting it within 0.2 points of Claude Opus 4.6. On Terminal-Bench 2...
The Segment co-founder published "Small models have arrived" on August 26, and it took 703 points on Hacker News (calv.info). His measurement: a personalized-news task that cost about a dollar on Sonnet-class models now runs at about a dime. Ten times cheaper, doing the job we...
Alibaba released Qwen3.6-27B on April 22. Dense architecture. Open weights. 77.2% on SWE-bench Verified, within 3.7 points of Claude Opus 4.6. On SkillsBench, it scores 48.2% versus its own 397B MoE predecessor's 30.0%. That's a 77% improvement with 14.8x fewer parameters. Let...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.