Fetching from the wire…
Public story · 2026-02-22 · source-backed
The biggest single-benchmark jump in a frontier model update: 37.6% → 68.8% on ARC-AGI-2. For comparison, GPT-5.2 scored 54.2%, Gemini 3 Pro 45.1%. ARC-AGI-2 measures novel problem-solving on adversarially constructed tasks. A near-doubling suggests genuine reasoning improvement, not benchmark optimization. ARC Prize
Each link below shares sources, entities, or timing with this story.
Google released Gemini 3.1 Pro on February 19, the first ".1" increment in Gemini's history. The standout metric: 77.1% on ARC-AGI-2, more than double the reasoning performance of Gemini 3 Pro. VentureBeat calls it "Deep Think Mini" — adjustable reasoning depth on demand. Feat...
GPT-5.4 scores 0.26%. Opus 4.6 scores 0.25%. Grok-4.20 scores 0.00%. Humans score 100%. The Decoder covered the ARC-AGI-3 launch on March 25, and the results make every "AGI is here" claim look premature. François Chollet launched ARC-AGI-3 at Y Combinator HQ alongside a fires...
Everyone spent yesterday arguing about benchmark numbers. Tencent quietly published data suggesting the numbers belong to your infrastructure, not the model. The WorkBuddy Bench leaderboard reports every model under two different agent harnesses — CodeBuddy Code and Claude Cod...
ARC-AGI-2 Reasoning Race. Gemini 3.1 Pro hit 77.1% and Claude Opus 4.6 hit 68.8% — both more than doubling their predecessors' scores in a single generation. The entire field moved ~10 points over the prior two years; each model jumped 30-40 points in one release. The top thre...
Four frontier models. Five sealed engineering problems. The result everybody will quote is that Claude Fable 5 won. The result that should actually change how you work is buried three-quarters down the page. JuliaHub published an evaluation on July 30 running four frontier mod...
43.3% on Frontier-Bench v0.1. Opus 4.8 scored 18.7%. That's not an incremental bump, that's the same benchmark with a different shape of answer. Anthropic released Claude Opus 5 on July 24 at $5/$25 per million input/output tokens, exactly half of Fable 5's $10/$50, while matc...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.