Fetching from the wire…
Top 5 · 2026-04-26 · source-backed
DeepSeek released V4 on April 24 and the numbers demand attention. V4-Pro is 1.6 trillion parameters total with 49 billion active, MIT-licensed, native 1M-token context. It scores 80.6% on SWE-bench Verified, putting it within 0.2 points of Claude Opus 4.6. On Terminal-Bench 2.0 it actually leads Claude: 67.9% vs 65.4%. LiveCodeBench: 93.5% vs 88.8%.
V4-Flash is the more interesting play for production use. 284B total parameters, 13B active, priced at $0.14/$0.28 per million tokens. That's the cheapest model in its performance tier by a wide margin. LMSYS benchmarked throughput at 199 tokens/sec on B200 and 266 tokens/sec on H200, with throughput dropping only 10% from 4K to 900K context tokens. That's a remarkably flat scaling curve.
The architecture is genuinely novel. Hybrid Attention combining Compressed Sparse Attention and Heavily Compressed Attention cuts inference FLOPs to 27% and KV cache to 10% compared to V3.2 at 1M context. Community analysis on r/LocalLLaMA shows KV cache dropping from 83.9 GiB to 9.62 GiB at 1M tokens. That's a 10x reduction, and it's what makes the pricing possible.
Then there's the chip story. Fortune reports Huawei confirmed V4 runs on Ascend 950 supernodes. First trillion-parameter MoE deployed without any NVIDIA hardware. DeepSeek rewrote their stack from CUDA to Huawei's CANN framework, with Cambricon and Moore Threads chips also supported. US export controls were supposed to prevent exactly this. They didn't.
What should you do? If you're running any batch processing, RAG pipelines, or high-volume agent workflows, V4-Flash deserves a serious evaluation this week. At $0.14/M input tokens versus $3/M for Claude Sonnet or $15/M for Opus, the cost difference funds a lot of quality-checking infrastructure. I'm not saying it replaces Opus for complex reasoning. I'm saying for 70% of the token volume in most production systems, it might be good enough at 1/100th the price. That's worth testing.
Each link below shares sources, entities, or timing with this story.
DeepSeek released V4 Preview on April 24 with two open-weight variants: V4-Pro (1.6T total parameters, 49B activated via MoE) and V4-Flash (284B parameters, 13B activated). Both support 1M-token context windows. Both are Apache 2.0 licensed. Both are live right now on Hugging...
Alibaba released Qwen3.6-27B on April 22. Dense architecture. Open weights. 77.2% on SWE-bench Verified, within 3.7 points of Claude Opus 4.6. On SkillsBench, it scores 48.2% versus its own 397B MoE predecessor's 30.0%. That's a 77% improvement with 14.8x fewer parameters. Let...
For about a year, "run your agent locally" meant accepting a model that couldn't reliably call a tool twice in a row. That excuse is gone. Meta Superintelligence Labs published Muse Glimmer today: a 29.6B dense causal transformer, 52 layers, 6,656 hidden dim, with a ~1.8B ViT-...
DeepSeek-V4-Flash-0731 landed July 31 under MIT with a DSpark speculative-decoding module attached. Terminal Bench 2.1: 82.7. Toolathlon-Verified: 70.3. DSBench-FullStack: 68.7. DeepSWE: 54.4. NL2Repo: 54.2. The model card claims it beats DeepSeek-V4-Pro (Preview) "despite its...
OpenAI shipped GPT-5.5 on April 23, six weeks after 5.4. The capability jump is real: 82.7% on Terminal-Bench 2.0 vs Claude Opus 4.7's 69.4%. The Pro tier nearly doubles Opus 4.7 on FrontierMath Tier 4 at 39.6% vs 22.9%. It uses 40% fewer tokens on Codex tasks while matching 5...
Martin Alderson's essay "The upcoming AI margin collapse, part 1: GLM 5.2" hit 675 points and 462 comments on Hacker News, and it's the rare HN chart-topper that's actually about spreadsheet math instead of vibes. The argument is simple. Z.ai's GLM 5.2 delivers frontier-adjace...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.