Fetching from the wire…
Public story · 2026-07-30 · high
23 frontier models summarized alerts they'd already gotten, but skipped the raw disk data where hidden intrusions live.
Why now: The benchmark posted July 30, arguing incident response, the exact workload agents are pitched for, is the one they're failing at.
Researchers tested 23 frontier AI models against a simulated network breach, per the new SecRespond benchmark, arXiv 2607.26791. Zero of the 23 closed out a single range end to end, a real gap for teams pitching AI agents at incident response.
Ten cyber ranges made up the test, covering four entry-point types, 21 ATT&CK techniques, and five operating systems, evaluated on the OpenCode harness.
Each agent gets a forensic disk snapshot of a breached host, plus alerts, vulnerability scans, and baseline security checks. It has to produce three reports, on intrusion, baseline risk, and vulnerability risk, plus a remediation plan.
Same failure mode across every model. Agents write solid summaries of what the alerts already caught, but skip the disk snapshot as a lead worth chasing on its own. Intrusions that never tripped an alert mostly go unfound.
I'd bet these models keep writing clean reports on what the alerts already flagged, and keep missing what they didn't. Explaining a known incident and hunting for an unknown one are different skills, and right now only the first one works.
Each link below shares sources, entities, or timing with this story.
The Pragmatic Engineer published a deep read on August 25 of Inspect, the coding agent Ramp built instead of standardizing on Claude Code or Cursor. The numbers: Inspect authors 75% of Ramp's merged PRs, 90% of PRs in its own repository, passed 1 million total sessions in July...
v0.10.0 (~84.8k stars, Apache-2.0) ships no agent of its own and drives whichever CLI you already have, Claude Code, Codex, Cursor, Copilot, OpenClaw, Gemini, Kimi, Qwen, Cline, plus BYOK OpenAI-compatible endpoints, via od mcp install <agent>. It produces single-page HTML pro...
$3,054 against $38,370. Same benchmark, better score. Praxist (arXiv 2608.25955, submitted August 26) replaces per-attempt agent memory with a typed evidence graph of findings, plus lane-structured frontiers and agendas, so later attempts inherit validated mechanisms rather th...
Everyone writing SKILL.md files has absorbed the same folklore. Keep the top file thin. Push detail into reference files. Let the agent walk the tree as needed. More layers, more context efficiency. A controlled study submitted July 20 tested that across InfiniteBench, three a...
A paper from Xiao Yu, Baolin Peng, and Ruize Xu makes a claim that seems obvious once stated and is genuinely new as a training methodology: modern agents are inseparable from their inference harnesses, so training them in stripped-down RL sandboxes produces a train/serve mism...
The antigravity-awesome-skills repository hit 36,145 GitHub stars with a catalog of 1,400+ installable SKILL.md playbooks that work across Claude Code, Cursor, Codex CLI, Gemini CLI, Kiro, OpenCode, and GitHub Copilot. One command: npx antigravity-awesome-skills --claude. That...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.