Fetching from the wire…
Public story · 2026-08-10 · high
The results are Anthropic's evidence for switching Claude Code's auto mode on by default for every user, per TechCrunch.
Why now: Anthropic released the eval alongside the news that Claude Code's auto mode is becoming the default, per TechCrunch.
Trajectory Labs threw 720 indirect prompt injection attacks at Claude's auto mode, and none broke through, per TechCrunch.
Anthropic is using the results to justify switching Claude Code's auto mode on by default, the outlet reported. That means people who never touched that setting will soon have Claude acting on its own without asking first.
The independent evaluator ran 72 attack scenarios, each ten times, against Fable 5, Opus 5, and Sonnet 5 as of July 17. Anthropic didn't design or see the scenarios in advance, which is a meaningfully different claim than a vendor grading its own homework.
That's real. A held-out third-party eval beats internal red-teaming. A clean sweep across three models and 720 attempts is a specific, checkable number, not a vibe, and I'll take that seriously. But 72 scenarios is a slice of the ways someone can hide instructions in a file, a page, or a tool result Claude reads mid-task. I don't think zero successes against a known test set tells you much about attacks nobody thought to write yet.
Watch whether Anthropic or Trajectory Labs publishes the scenario list, or repeats the test after auto mode ships to everyone. A single clean run before launch is a snapshot, not a track record.
Anthropic released the eval alongside the news that auto mode is becoming Claude Code's default, per TechCrunch.
Each link below shares sources, entities, or timing with this story.
43.3% on Frontier-Bench v0.1. Opus 4.8 scored 18.7%. That's not an incremental bump, that's the same benchmark with a different shape of answer. Anthropic released Claude Opus 5 on July 24 at $5/$25 per million input/output tokens, exactly half of Fable 5's $10/$50, while matc...
This one's been building for days and it crystallized this week. Per The Register, the incident behind the US export-control block on Anthropic's Fable 5 and Mythos 5 wasn't a jailbreak or a guardrail bypass. It was a plain three-word prompt, "fix this code," run against CVE-l...
This one changed how I'm spending my week. Anthropic's July 24 context-engineering post says they removed over 80% of Claude Code's system prompt for Opus 5 and Fable 5 with no measurable loss on coding evals. They call it "unhobbling" — stripping guardrails and rules that new...
The extension takes autonomous actions (reading, clicking, form-filling) without per-action approval, gated by a safety classifier validating each action plus probes scanning page content for injection attempts. Anthropic publishes per-model attack success rates with safeguard...
Two data points that tell the same story. First, Value Add Pulse counts four frontier launches in 30 days: Gemini 3.5 Pro, Grok 5, Anthropic's Fable 5 and Mythos 5, plus open-weight GLM-5.2 and Kimi K2.7. The model-layer moat compressed from quarters to weeks. Second, TechCrun...
TechCrunch reports the July 23 macOS and Windows rollout, built on ChatGPT-Live. It drives ChatGPT Work and Codex, so you can direct several agents simultaneously by voice, and macOS "Appshots" let it read the frontmost window including alt-text. Global rollout to Plus, Pro, B...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.