Fetching from the wire…
Public story · 2026-08-25 · high
Sonnet 5 recalled Rails APIs at 25.4%, nearly double Opus 4.8's 15.9%, on Rails' new open benchmark.
Why now: The scores come from Rails' August 24 release of lemans, the first outside look at how frontier models handle a real Rails codebase.
Rails released lemans, an open-source agent harness for Ruby, on August 24, then used it to benchmark four coding agents on 63 real Rails tasks.
The core team built it because Ruby developers shouldn't need Harbor, the existing Python-based harness, to grade an agent on Rails code. The 63-task suite is the first public scorecard for how four models handle real Rails code, not generic web work.
ox-alpha scored 52 of 63 tasks, the best result among the four models tested.
Terra scored 49 of 63, at $0.20 a run and a 182-second median.
Qwen 3.8-27B, the open-weight entry, scored 48 of 63. Its median run took 27 minutes, long enough that Rails raised the task timeout from 30 minutes to 60.
Sonnet 5 scored 44 of 63, the weakest result Anthropic has recorded on the benchmark. It recalled Rails APIs more often than Opus 4.8, 25.4% against 15.9%.
The release doesn't say whether Rails plans to re-run the suite as new models ship.
Each link below shares sources, entities, or timing with this story.
Elvis Saravia's roundup characterizes it as making plans, driving browsers and terminals, and finishing multi-step tasks where prior Sonnets stopped short, notably verifying its own output unprompted. Anthropic puts it near Opus 4.8 on reasoning, tool use, coding, and knowledg...
This is the most useful thing I read this week and it isn't close. Anthropic published its internal methodology for running large-scale code migrations with Claude Code on July 16, and unlike most engineering-blog playbooks, it carries receipts. Bun's Zig→Rust migration: rough...
Every conversation I've had about AI costs in the last six months eventually lands on the same tension: you want the smartest model for the hard decisions, but you can't afford to run it on every token. Anthropic just gave that tension a formal solution. The advisor tool, now...
mvp-game-loop. Eighteen out of thirty agents, independently, chose the identical branch name. Anthropic's Frontier Red Team published the first large study of multiagent failure modes on August 13, testing Sonnet 4.6/5, Opus 4.6/4.8, and Mythos Preview/5 across four categories...
The extension takes autonomous actions (reading, clicking, form-filling) without per-action approval, gated by a safety classifier validating each action plus probes scanning page content for injection attempts. Anthropic publishes per-model attack success rates with safeguard...
Anthropic commissioned the independent evaluator to test 72 injection scenarios, held out from Anthropic, each run 10 times against Fable 5, Opus 5, and Sonnet 5 as of July 17. Clean sweep. TechCrunch has the details. A third-party held-out eval is a much stronger claim than i...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.