Fetching from the wire…
Public story · 2026-08-04 · high
At Q3 it edges out Qwen3.6-27B on big coding jobs; drop to Q2 and Qwen wins, per a 189-upvote r/LocalLLaMA thread.
Why now: The thread is framed as reactions after the first weekend of hands-on use with the 0731 build, making this the earliest wide read on how it holds up quantized.
V4-Flash-0731 loses far more accuracy under quantization than DeepSeek's preview build, per a 189-upvote r/LocalLLaMA thread.
That matters because the model's standing against Qwen3.6-27B flips depending on the quant level someone actually runs.
Posters in the thread describe Q2 and Q3 builds as functionally different models, with reasoning traces shaped unlike the full-precision version's own.
The thresholds are specific. At Q3, V4-Flash-0731 finally out-codes Qwen3.6-27B on large repositories and in harnesses running system prompts past 30,000 tokens. Drop to Q2, and Qwen3.6-27B at Q8 wins outright.
Unsloth's KL-divergence charts posted in the thread back up the complaint. The preview build stayed close to its full-precision outputs even at aggressive quant levels. The 0731 release shows poor KL-divergence even at IQ4_XS, a milder quant level than the Q2 and Q3 builds drawing complaints.
The thread doesn't say whether DeepSeek changed its training process between the preview and 0731, only that the quantization behavior changed.
Each link below shares sources, entities, or timing with this story.
DeepSeek released V4 on April 24 and the numbers demand attention. V4-Pro is 1.6 trillion parameters total with 49 billion active, MIT-licensed, native 1M-token context. It scores 80.6% on SWE-bench Verified, putting it within 0.2 points of Claude Opus 4.6. On Terminal-Bench 2...
The aggregate parameter count Chinese labs shipped in 30 days now exceeds what the whole open-weight ecosystem produced in the first half of 2026. r/LocalLLaMA Nikkei Asia reported August 15 that Z.ai positions GLM-5.3 as a direct rival to Anthropic's Mythos on coding and secu...
DeepSeek released V4 Preview on April 24 with two open-weight variants: V4-Pro (1.6T total parameters, 49B activated via MoE) and V4-Flash (284B parameters, 13B activated). Both support 1M-token context windows. Both are Apache 2.0 licensed. Both are live right now on Hugging...
The leaderboard says first place. The methodology says you should check your own bill. Qwen3.8 Max now ranks first on Artificial Analysis' agentic index, scoring 86.1 on OSWorld-Verified ahead of GPT-5.6 Sol Max at 83.2 and Fable 5 at 85.0, priced at $2.00/M input and $6.00/M...
For about a year, "run your agent locally" meant accepting a model that couldn't reliably call a tool twice in a row. That excuse is gone. Meta Superintelligence Labs published Muse Glimmer today: a 29.6B dense causal transformer, 52 layers, 6,656 hidden dim, with a ~1.8B ViT-...
The repo appeared August 24, opening the weights of a multimodal MoE that Qwen frames explicitly as an architecture preview, the same role Qwen3-Next played for Qwen3.5 (GitHub). The hybrid Gated DeltaNet plus Gated Attention design it previews already carried through the Qwen...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.