Fetching from the wire…
Public story · 2026-02-22 · source-backed
Largest open-source sparse MoE model: 400B parameters, 13B active per token, trained on 17 trillion tokens. Novel SMEBU for load balancing. Zero loss spikes with Muon optimizer across 17T tokens — a remarkable training stability achievement. arXiv | Arcee Blog
Each link below shares sources, entities, or timing with this story.
ARC-AGI-2 Reasoning Race. Gemini 3.1 Pro hit 77.1% and Claude Opus 4.6 hit 68.8% — both more than doubling their predecessors' scores in a single generation. The entire field moved ~10 points over the prior two years; each model jumped 30-40 points in one release. The top thre...
36. arXiv — CUWM 37. arXiv — IntentCUA 38. arXiv — SpargeAttention2 39. arXiv — Arcee Trinity Large 40. ARC Prize — ARC-AGI-2 41. Arcee Blog ---
The architecture report describes a 125B sparse MoE activating 6B parameters per token, plus 51B of n-gram embedding tables held off the accelerator in host memory with prefetching (arXiv 2608.30320). Against the prior 397B-A17B model it leads on 8 of 14 pre-training benchmark...
Simon Willison found it in the OpenRouter price list. 1.6T MoE, ~49B active, 1M context, up to 384K output, at $0.435/M input (cache miss) and $0.87/M output. The agentic-coding deltas versus the preview are the story: DeepSWE 12.8 → 62.7, CyberGym 52.7 → 83.3, Terminal Bench...
Released August 6 by InclusionAI, it's a native hybrid-reasoning MoE with switchable Thinking and Instant modes, native function calling, prompt caching and 256K context, aimed at resource-constrained and fully local deployment (demoed against obsidian-cli over a local notes r...
GLM-5.1 from Zhipu AI scored 58.4% on SWE-bench Pro. GPT-5.4 scored 57.7%. Claude Opus 4.6 scored 57.3%. That's the first time an open-weight model has ever topped a major coding benchmark against the best proprietary models. The specs matter. GLM-5.1 is a 754B-parameter mixtu...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.