Fetching from the wire…
Top 5 · 2026-03-29 · source-backed
GPT-5.4 scores 0.26%. Opus 4.6 scores 0.25%. Grok-4.20 scores 0.00%. Humans score 100%. The Decoder covered the ARC-AGI-3 launch on March 25, and the results make every "AGI is here" claim look premature.
François Chollet launched ARC-AGI-3 at Y Combinator HQ alongside a fireside conversation with Sam Altman. The benchmark is fully interactive: hundreds of game-style environments with no instructions, no rules, and no stated goals. Agents must figure out what to do by exploring and learning from environmental feedback. Sustained sequential reasoning, state tracking across hundreds of steps, real-time adaptation. Everything current language models can't do.
Here's what stopped me cold: simple CNN and graph-search approaches scored 12.58%. Over 30x better than any frontier LLM. Not a fine-tuned model, not a billion-dollar training run. Basic pattern matching over a narrow domain demolishes trillion-parameter models on tasks requiring actual novel reasoning. The models interpolate beautifully from training data. They can't extrapolate at all. The gap between "looks smart" and "is smart" has never been measured this precisely.
The Chollet-Altman pairing is notable. They agreed on a timeline: AGI "probably by early 2030s, around ARC-AGI 6 or 7" (OfficeChai). The creator of the hardest AI benchmark and the CEO of the company most invested in scaling sat together and acknowledged that current approaches aren't sufficient. Meanwhile, GPT-5.4 saturated USAMO 2026 at 95%, up from near-zero last year (MathArena). So models are getting dramatically better at pattern-matchable math competitions while scoring essentially zero on tasks requiring genuine novel reasoning. That divergence is the whole story.
ARC Prize 2026 offers $2M+ in prizes. All solutions must be open-sourced. If you want to work on the hardest unsolved problem in AI, the benchmark and the funding are waiting. This connects directly to the Anthropic harness story: if individual models can't reason sequentially, you build systems where reasoning is distributed across specialized agents with external feedback loops. Multi-agent architectures aren't just a design pattern. They're a workaround for a fundamental limitation.
Each link below shares sources, entities, or timing with this story.
The ARC Prize Foundation dropped ARC-AGI-3 on March 25 and the results broke my mental model of how AI capability scales. Symbolica's Arcgentica framework scored 36.08% (113 of 182 playable levels, 7 of 25 games completed) using Claude Opus 4.6 as its backbone. Cost: $1,005. F...
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's...
GLM-5.1 from Zhipu AI scored 58.4% on SWE-bench Pro. GPT-5.4 scored 57.7%. Claude Opus 4.6 scored 57.3%. That's the first time an open-weight model has ever topped a major coding benchmark against the best proprietary models. The specs matter. GLM-5.1 is a 754B-parameter mixtu...
Microsoft announced Critique on March 30. Here's how it works: when you use M365 Copilot Researcher, GPT drafts the initial research response. Then Claude reviews it for accuracy, completeness, and citation quality. You only see the final result after both models have had thei...
Everyone spent yesterday arguing about benchmark numbers. Tencent quietly published data suggesting the numbers belong to your infrastructure, not the model. The WorkBuddy Bench leaderboard reports every model under two different agent harnesses — CodeBuddy Code and Claude Cod...
19. Karpathy — microGPT 20. TechCrunch — Altman vs Anthropic 21. MIT Tech Review — LeCun AMI Labs 22. Dario Amodei — Adolescence essay 23. Simon Willison — Showboat and Rodney 24. Lenny's Newsletter — v0 25. ARC Prize — ARC-AGI-3
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.