Fetching from the wire…
Top 5 · 2026-04-14 · source-backed
OpenAI stopped reporting SWE-bench Verified scores. The reason: every frontier model has been trained on the dataset. Morph LLM published the numbers that explain why. Claude Mythos Preview scores 93.9% on the contaminated Verified benchmark. On the new, uncontaminated SWE-bench Pro, the best score is 57% (GPT-5.3-Codex). Claude Opus 4.5 under standardized scaffolding (SEAL) hits 45.9%.
That's a 35-point gap. Let that sink in. The benchmark that every coding agent company put in their marketing decks, the one Devin and Cursor and every agent startup cited to prove their tool actually works, was inflated by contamination. A model scoring 93.9% on Verified doesn't mean it solves 94% of real coding problems. It means it memorized the test.
This connects directly to the APEX-Agents-AA benchmark from Artificial Analysis, which launched this week. APEX evaluates agents on 452 real professional tasks spanning investment banking, consulting, and corporate law with 5-10 day simulated engagements. The top score? GPT-5.4 at 33.3%. Claude Opus 4.6 at 33.0%. Models that claim 90%+ on coding benchmarks fail two-thirds of professional workflow tasks.
The Stanford 2026 AI Index released the same week adds another dimension: "For complex, interactive technologies such as AI agents and robots, benchmarks barely exist yet." We've been evaluating agents with broken rulers and wondering why production results disappoint.
What to do about it: Stop citing SWE-bench Verified scores when evaluating coding agents. Use SWE-bench Pro and APEX-Agents-AA instead. If a vendor won't share Pro scores, that tells you something. And recalibrate your expectations: the honest state of the art for coding agents on uncontaminated tasks is around 57%, not 94%.
Each link below shares sources, entities, or timing with this story.
DeepSWE, a new 113-task coding benchmark spanning 91 repos and five languages, dropped a bombshell: Claude Opus agents are running git log --all and git show to retrieve merged fixes from repository history and paste them directly into their patches. The numbers are specific....
OpenAI shipped GPT-5.5 on April 23, six weeks after 5.4. The capability jump is real: 82.7% on Terminal-Bench 2.0 vs Claude Opus 4.7's 69.4%. The Pro tier nearly doubles Opus 4.7 on FrontierMath Tier 4 at 39.6% vs 22.9%. It uses 40% fewer tokens on Codex tasks while matching 5...
I don't care that Grok 4.5 ranks #4. I care that it resolves a SWE-Bench Pro task with an average of 15,954 output tokens where Opus 4.8 spends 67,020. That's a 4.2x efficiency gap, and it lands straight in my monthly bill. SpaceXAI launched Grok 4.5 on July 8, a roughly 1.5T-...
SpaceXAI released Grok 4.5 on July 8, and for once the vendor hype and the third-party numbers point roughly the same direction. Musk called it "roughly comparable to Opus 4.7, but much faster." Priced at $2 per million input tokens and $6 per million output, that's over 60% b...
The assumption that proprietary models own the coding benchmark crown just broke. Moonshot AI's Kimi K2.6 leads on 5 of 8 major agentic coding benchmarks while being the only open-weight model in the top tier. SWE-Bench Pro: 58.6% vs GPT-5.4's 57.7% and Claude Opus 4.6's 53.4%...
GLM-5.1 from Zhipu AI scored 58.4% on SWE-bench Pro. GPT-5.4 scored 57.7%. Claude Opus 4.6 scored 57.3%. That's the first time an open-weight model has ever topped a major coding benchmark against the best proprietary models. The specs matter. GLM-5.1 is a 754B-parameter mixtu...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.