Fetching from the wire…
Public story · 2026-03-08 · source-backed
537 test cases across 8 categories testing 6 commercial products. Composite scores range from ~39 to ~98. Critical finding: providers catching >95% of prompt injections miss most unauthorized tool calls — tool abuse detection is universally weak. Provenance verification is nearly absent. First empirical comparison of the agent security market. GitHub
Each link below shares sources, entities, or timing with this story.
Two competing models for AI-powered security shipped on the same day. OpenAI launched Codex Security ("Aardvark") — an AI AppSec agent that builds project-specific threat models, then hunts for vulnerabilities and tests them in isolated environments. 30-day beta: 1.2M+ commits...
First standardized benchmark for AI agent skills. 86 tasks, 11 domains, 7,308 test trajectories. Critical finding: curated skills +16.2%, self-generated skills +0%. Run it against your own skills. GitHub | Paper
GitHub | Go, MIT Extremely new but architecturally significant. Performs static analysis on skill files (markdown, YAML, JSON) to detect threats before deployment — offline, deterministic, no LLM required. 138 detection rules across 15 categories. Companion service scanned 31,...
Released today, v0.21.13 rebuilds the /review skill around dedicated platform subcommands (meta, fetch-diff) instead of raw gh commands issued through prompt prose, and caps posted suggestions to Critical findings after round 5 to stop review loops (GitHub). Operationally it a...
affaan-m/ECC (36.3k forks, MIT) bundles 67 agents, 284 skills, 94 legacy command shims, and "instincts", patterns learned from prior sessions with confidence scores that auto-recall when relevant, plus a .ecc/memory/ markdown vault that's explicitly cross-harness, so context s...
ECC (238,911 stars, pushed August 9) shipped v2.1 with Plan Canvas, a browser review surface where you click-annotate a plan and approve or request changes instead of retyping corrections into the terminal, plus a native install target for Moonshot's Kimi Code and an optional...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.