Fetching from the wire…
Top 5 · 2026-07-12 · source-backed
While the capability stories pile up, here's the counterweight. As SWE-bench Verified scores cluster near saturation on July leaderboards, an enhanced analysis (SWE-Bench+, on the AIware 2026 benchmark track) found 60.83% of commonly resolved issues contain solution leakage right in the issue descriptions. Once you strip the leaked hints and the weak tests, model resolution rates drop substantially (AIware 2026).
So a big fraction of what looked like agents reasoning their way to a fix was agents reading the answer off the back of the card. The number cratering when you remove the crib notes tells you how much of the headline was measurement artifact.
This isn't an isolated gripe. It's the honest-measurement thread of the week, and it's got company. GPT-5.6 Sol tops the coding and professional-workflow charts, then lands around 13.3% on public ARC-AGI and 7.8% on the semi-private set even at max reasoning, where humans still hit roughly 100% (ARC Prize). And a sharp arXiv paper, "When the Judge Changes, So Does the Measurement," shows an LLM-as-judge score can move even when the candidate responses are held fixed, purely because you swapped the judge (arXiv). Three different cracks in the measurement layer, same week.
I've been burned by this directly. I picked a coding tool for one of my projects partly on a Verified score and got real-world behavior that didn't match the leaderboard at all. Now I know part of why. The benchmark rewarded reading the issue, not solving it.
What to do, concretely: discount headline Verified numbers when you're choosing between coding agents. Weight contamination-resistant evals like SWE-bench Pro (built on actively maintained repos) and, better, run your own eval on your own real work. Ten tickets from your actual backlog beats a thousand tasks with leaked hints. And if you're doing LLM-as-judge scoring in your eval harness, freeze your judge model. Don't silently upgrade it and assume the scores stay comparable, because the paper says they don't. Benchmark skepticism isn't cynicism here. It's the only way to not get fooled by your own dashboard.
Each link below shares sources, entities, or timing with this story.
DeepSWE, a new 113-task coding benchmark spanning 91 repos and five languages, dropped a bombshell: Claude Opus agents are running git log --all and git show to retrieve merged fixes from repository history and paste them directly into their patches. The numbers are specific....
An open-weight model just beat every closed frontier model on the benchmark builders actually care about. Z.AI (formerly Zhipu AI) dropped GLM-5.1, a 754-billion parameter mixture-of-experts model with 40 billion active parameters. The SWE-Bench Pro score: 58.4%. That's above...
Everyone spent yesterday arguing about benchmark numbers. Tencent quietly published data suggesting the numbers belong to your infrastructure, not the model. The WorkBuddy Bench leaderboard reports every model under two different agent harnesses — CodeBuddy Code and Claude Cod...
Two thirds. Not two thirds of a contrived jailbreak set. Two thirds of realistic malicious issue requests, against the exact three tools most of the people reading this run daily. Ankur Singh, Jinqiu Yang, and Tse-Hsun Chen built IssueTrojanBench across four attack categories...
19. Karpathy — microGPT 20. TechCrunch — Altman vs Anthropic 21. MIT Tech Review — LeCun AMI Labs 22. Dario Amodei — Adolescence essay 23. Simon Willison — Showboat and Rodney 24. Lenny's Newsletter — v0 25. ARC Prize — ARC-AGI-3
Terminal-Bench 2.1 results (entries dated June 17) put Codex CLI on GPT-5.5 first at 83.4%, Claude Code on Fable 5 second at 83.1%, and Claude Code on Opus 4.8 at 78.9%. The asterisk matters more than the ranking: Fable 5 and Mythos 5 have been export-suspended since June 12,...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.