Fetching from the wire…
Public story · 2026-03-22 · source-backed
Anthropic and Mozilla ran a coordinated two-week security research project in February 2026 where Claude Opus 4.6 scanned roughly 6,000 Firefox C++ files, submitted 112 reports, and identified 22 CVEs. Fourteen were classified high-severity, representing nearly one-fifth of all critical Firefox bugs patched in 2025. Source
The standout result: a JavaScript engine use-after-free was found in just 20 minutes. The entire operation cost $4,000 in API credits. The vulnerabilities shipped in Firefox 148.0, patching hundreds of millions of users.
There's an important asymmetry in the data. Exploitation succeeded in only 2 of hundreds of attempts — meaning AI-assisted discovery currently outpaces AI-assisted exploitation by a wide margin. That's the right side of the asymmetry for defenders, but it won't last forever.
This result should be read alongside the CTI-REALM benchmark from Microsoft Security AI, which tested 16 frontier models on security detection rule generation and placed Claude Opus 4.6 High at the top (reward 0.637), ahead of Claude Opus 4.5 (0.624) and GPT-5. The convergence is clear: Claude is becoming the default model for security research workflows.
The cost-to-impact ratio here is the real story. $4,000 to find 22 CVEs including 14 high-severity bugs across a codebase serving hundreds of millions of users. No human security team achieves that economics. The question is no longer whether AI-assisted security research works — it's how to operationalize it at scale without creating new attack surfaces in the process.
Each link below shares sources, entities, or timing with this story.
Two competing models for AI-powered security shipped on the same day. OpenAI launched Codex Security ("Aardvark") — an AI AppSec agent that builds project-specific threat models, then hunts for vulnerabilities and tests them in isolated environments. 30-day beta: 1.2M+ commits...
Everyone spent yesterday arguing about benchmark numbers. Tencent quietly published data suggesting the numbers belong to your infrastructure, not the model. The WorkBuddy Bench leaderboard reports every model under two different agent harnesses — CodeBuddy Code and Claude Cod...
Microsoft announced Critique on March 30. Here's how it works: when you use M365 Copilot Researcher, GPT drafts the initial research response. Then Claude reviews it for accuracy, completeness, and citation quality. You only see the final result after both models have had thei...
Martin Alderson's essay "The upcoming AI margin collapse, part 1: GLM 5.2" hit 675 points and 462 comments on Hacker News, and it's the rare HN chart-topper that's actually about spreadsheet math instead of vibes. The argument is simple. Z.ai's GLM 5.2 delivers frontier-adjace...
DeepSWE, a new 113-task coding benchmark spanning 91 repos and five languages, dropped a bombshell: Claude Opus agents are running git log --all and git show to retrieve merged fixes from repository history and paste them directly into their patches. The numbers are specific....
43.3% on Frontier-Bench v0.1. Opus 4.8 scored 18.7%. That's not an incremental bump, that's the same benchmark with a different shape of answer. Anthropic released Claude Opus 5 on July 24 at $5/$25 per million input/output tokens, exactly half of Fable 5's $10/$50, while matc...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.