Fetching from the wire…
Public story · 2026-07-27 · high
Unprotected, Opus 5 still fails 3.70% of browser injection tests, per the system card; Sonnet 5 alone scored 0.93%.
Why now: Simon Willison's July 25 post spotlighting Boris Cherny's page-73 claim is what's putting Opus 5's injection numbers in front of builders.
Opus 5 blocked every attempt in 129 browser prompt-injection tests when paired with Auto Mode's two-layer defense, per its system card. That's the number agent builders want: a hidden instruction buried in a page has been the easiest way to hijack a browsing agent.
The model alone doesn't get there. Its unprotected success rate was 3.70%, down from 31.5% in Anthropic's own testing. Auto Mode adds two checks: one layer scans incoming page data for hidden instructions, a second blocks dangerous actions before they run.
Gray Swan's independent test found a similar drop: Opus 5 failed 2.0% after 15 attempts, down from 5.5% for Opus 4.8. Sonnet 5 beat Opus 5's unprotected score anyway, coming in at 0.93% with no product-side defenses running at all.
Simon Willison surfaced Boris Cherny's claim that Opus 5 is Anthropic's least injectable model, citing page 73 of the system card. He chose to surface the number rather than pick it apart, and that choice is the signal.
Each link below shares sources, entities, or timing with this story.
Boris Cherny told Simon Willison the most exciting thing about Opus 5 isn't the evals, it's the injection resistance, and that it's "a bit buried in the system card" on page 73. The numbers: attacker success within 15 attempts on the Gray Swan indirect injection benchmark fell...
Willison documented on July 4 that Opus 4.8 and Sonnet 5 can perform worse than older versions when driving bespoke file-edit tools, because they're increasingly trained and optimized for Claude Code's native editor format (Simon Willison). This is a real trap if you're buildi...
No new posts today. Willison's Feb 17-21 output was extraordinary: 10+ posts covering Sonnet 4.6, GGML/HuggingFace merger, SWE-bench analysis (Opus 4.5 leads at 76.8%, Chinese models dominate top 10), and the Karpathy "Claws" amplification. His Showboat ecosystem — Rodney, Cha...
Anthropic shipped Fable 5 on June 9. Willison spent ~5.5 hours stress-testing it: slow and expensive, but it handled everything he threw at it, including agentic coding. (Simon Willison) The tell that it's a real working model and not a benchmark queen: because it post-dated A...
Every conversation I've had about AI costs in the last six months eventually lands on the same tension: you want the smartest model for the hard decisions, but you can't afford to run it on every token. Anthropic just gave that tension a formal solution. The advisor tool, now...
If you've used Claude Code for any serious session, you know the drill. Approve. Approve. Approve. Approve. You stop reading the prompts after the fifteenth one. That's the worst possible security outcome, way worse than a well-designed automated check. Anthropic launched auto...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.