Fetching from the wire…
Public story · 2026-03-08 · source-backed
The first benchmark built on continuous integration loops evaluates agents on long-horizon codebase maintenance (average 233 days, 71 consecutive commits per task). Most models achieve a zero-regression rate below 0.25 — only Claude Opus exceeds 0.5. Even frontier agents that pass initial patches introduce regressions later. This validates the gap between "write a fix" and "maintain a codebase." arXiv
Each link below shares sources, entities, or timing with this story.
Gergely Orosz published the first serious look at what AI coding actually costs at scale, and the numbers are wild. The Pragmatic Engineer covers "tokenmaxxing," a trend where engineers compete on AI token consumption leaderboards. At Meta, one engineer averaged 281 billion to...
The coding agent wars just entered a new phase. Cursor isn't just an IDE anymore. It's a model company. Cursor released Composer 2.5 on May 18 with a custom agentic coding model trained using 25x more synthetic tasks than Composer 2 and a novel "targeted textual feedback" appr...
A single Rust binary is saving agentic coding users 60-90% on token costs, and it took about five minutes to set up. RTK (Rust Token Killer) released v0.37.1 on April 18 and sits at 30,500 GitHub stars. The tool acts as a CLI proxy between your AI coding assistant and shell co...
OpenAI stopped reporting SWE-bench Verified scores. The reason: every frontier model has been trained on the dataset. Morph LLM published the numbers that explain why. Claude Mythos Preview scores 93.9% on the contaminated Verified benchmark. On the new, uncontaminated SWE-ben...
Cursor stopped being an IDE wrapper and became a model company. Cursor shipped Composer 2, a proprietary coding model trained via reinforcement learning on long-horizon coding tasks. On CursorBench — their own benchmark, caveats acknowledged — it scores 61.3, beating Claude Op...
A paper from Xiao Yu, Baolin Peng, and Ruize Xu makes a claim that seems obvious once stated and is genuinely new as a training methodology: modern agents are inseparable from their inference harnesses, so training them in stripped-down RL sandboxes produces a train/serve mism...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.