Fetching from the wire…
Public story · 2026-02-24 · source-backed
simonwillison.net — Four models now above 80% SWE-bench (Opus 4.6 at 80.8%, Gemini 3.1 Pro at 80.6%, GPT-5.3-Codex at 80.0%, Kimi K2.5 at 76.8%). Willison's analysis: we're approaching the ceiling where SWE-bench stops being a meaningful differentiator. The gap between models is now smaller than the gap between scaffolding approaches. This is the quantitative backing for the "scaffolding > model" thesis.
Each link below shares sources, entities, or timing with this story.
The February rankings reshuffled: Windsurf #1 (Arena Mode for side-by-side model comparison), Antigravity (Google) #2, Cursor #3 (8 async subagents + Multi-Agent Judging), Kimi Code NEW at #4 — the first open-source tool in the top 5 with 100-agent swarm capability backed by K...
OpenAI shipped GPT-5.5 on April 23, six weeks after 5.4. The capability jump is real: 82.7% on Terminal-Bench 2.0 vs Claude Opus 4.7's 69.4%. The Pro tier nearly doubles Opus 4.7 on FrontierMath Tier 4 at 39.6% vs 22.9%. It uses 40% fewer tokens on Codex tasks while matching 5...
OpenAI shipped the first model family explicitly designed for subagent pipelines. GPT-5.4 mini features a 400K context window, scores 54.4% on SWE-Bench Pro (vs. the flagship's 57.7%), and handles computer use at 72.1% on OSWorld — at $0.75 input / $4.50 output per million tok...
Two data points that tell the same story. First, Value Add Pulse counts four frontier launches in 30 days: Gemini 3.5 Pro, Grok 5, Anthropic's Fable 5 and Mythos 5, plus open-weight GLM-5.2 and Kimi K2.7. The model-layer moat compressed from quarters to weeks. Second, TechCrun...
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's...
An open-weight model just beat every closed frontier model on the benchmark builders actually care about. Z.AI (formerly Zhipu AI) dropped GLM-5.1, a 754-billion parameter mixture-of-experts model with 40 billion active parameters. The SWE-Bench Pro score: 58.4%. That's above...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.