Fetching from the wire…
Top 5 · 2026-03-12 · source-backed
OpenAI shipped native computer use in GPT-5.4, scoring 75.0% on OSWorld-Verified vs. the 72.4% human baseline (up from 47.3% in GPT-5.2). This is the first general-purpose model to surpass human performance on real desktop workflows. With 1M token context, 92.8% GPQA Diamond, and 83.3% ARC-AGI-2, frontier convergence is real — GPT-5.4, Opus 4.6, and Gemini 3.1 Pro now score within 2-3 points on most evaluations. The era of task-matched model selection has arrived. OpenAI Blog
Each link below shares sources, entities, or timing with this story.
OpenAI shipped the first model family explicitly designed for subagent pipelines. GPT-5.4 mini features a 400K context window, scores 54.4% on SWE-Bench Pro (vs. the flagship's 57.7%), and handles computer use at 72.1% on OSWorld — at $0.75 input / $4.50 output per million tok...
OpenAI released GPT-5.4 in Standard, Thinking, and Pro variants. Headline capabilities: native computer-use (75.0% on OSWorld-Verified, surpassing human 72.4%), 1M token context, and first-ever "compaction" support for longer agent trajectories. The Tool Search API is the buil...
Poetiq (6-person ex-DeepMind startup) hit 54% on ARC-AGI-2 at $30.57/task, beating Google's Gemini 3 Deep Think (45%, $77.16/task) through iterative refinement loops — no fine-tuning required. Meanwhile, Symbolica's Agentica reached 85.28% using recursive sub-agent delegation...
Google released Gemini 3.1 Pro on February 19, the first ".1" increment in Gemini's history. The standout metric: 77.1% on ARC-AGI-2, more than double the reasoning performance of Gemini 3 Pro. VentureBeat calls it "Deep Think Mini" — adjustable reasoning depth on demand. Feat...
GPT-5.4 scores 0.26%. Opus 4.6 scores 0.25%. Grok-4.20 scores 0.00%. Humans score 100%. The Decoder covered the ARC-AGI-3 launch on March 25, and the results make every "AGI is here" claim look premature. François Chollet launched ARC-AGI-3 at Y Combinator HQ alongside a fires...
Four frontier models. Five sealed engineering problems. The result everybody will quote is that Claude Fable 5 won. The result that should actually change how you work is buried three-quarters down the page. JuliaHub published an evaluation on July 30 running four frontier mod...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.