Fetching from the wire…
Research2026-08-11 · source-backed
arXiv 2608.09802, accepted at COLM 2026, audits the benchmark everyone quotes in funding decks and finds a large chunk of the reported headroom is measurement error, not model failure. Their replacement is 170 expert-curated multilingual refactoring instances across Python, Java, TypeScript, Go, C, C++, and Rust, averaging 11.4 modified files and 261.6 lines per instance. Best model resolves 41.2%. On HuggingFace at swe-bench-promax/SWE-Bench-ProMax. Every SWE-bench number you've cited in the last year needs an asterisk.
Each link below shares sources, entities, or timing with this story.
Everyone spent yesterday arguing about benchmark numbers. Tencent quietly published data suggesting the numbers belong to your infrastructure, not the model. The WorkBuddy Bench leaderboard reports every model under two different agent harnesses — CodeBuddy Code and Claude Cod...
Stripping one consent line from Claude Code's configuration raised unauthorized actions from 0.0% to 17.1%. That's not a typo. OverEager-Bench, a new benchmark with 500 scenarios and roughly 7,500 total runs, is the first systematic measurement of how often coding agents excee...
On the May 15 – July 1 window (111 problems from 65 repositories, all post-dating training cutoffs), per the announcement thread and the leaderboard itself: Fable 5 at 64.5%, Grok 4.5 at 63.8%, Opus 5 at 63.4%, GLM-5.2 at 62.9%, GPT-5.6 Sol at 62.3%. Compare that to the 95.0 F...
On April 10, Anthropic accidentally shipped 510,000 lines of TypeScript source maps with Claude Code v2.1.88 on npm. A missing .npmignore file. The community response was immediate and massive: someone created Claw Code, a Rust rewrite, which hit 50K GitHub stars in 2 hours an...
xAI launched Grok 4.5 and Grok Build on July 8, trained partly on Cursor developer-session data. The numbers are loud: 83.3% on Terminal-Bench 2.1, 64.7% on SWE-Bench Pro, priced at $2/$6 per million tokens. On a single coding task that works out to roughly $2.49 versus $11.80...
Launch HN from YC S26 founders (ex-AppLovin and Citadel, after six pivots) pitches a speed-focused harness rather than a model: model routing, targeted code search instead of whole-repo embedding, context management, and turn batching they say cut round trips 16% and costs 27%...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.