Fetching from the wire…
Public story · 2026-02-28 · source-backed
"Towards a Science of AI Agent Reliability" (arXiv 2602.16666) — 12 concrete metrics decomposing reliability along consistency, robustness, predictability, and safety. Key finding: stronger performance on benchmarks does NOT correlate with reliable real-world operation. Interactive dashboard included. (arXiv)
"Black-Box Reliability Certification for AI Agents" (arXiv 2602.21368) — A single reliability number per system-task pair via self-consistency sampling and conformal calibration with distribution-free guarantees. Directly implementable deployment gate for production agents. (arXiv)
"Prompt Injection Attacks on Agentic Coding Assistants" (arXiv 2601.17548) — First SoK across 78 studies. Attack success rates exceed 85% against SOTA defenses. 42 distinct attack techniques cataloged. Architectural mitigations required — filtering is fundamentally insufficient. (arXiv)
SUSVIBES Benchmark (arXiv 2512.03262) — 200 real-world tasks. SWE-Agent with Claude 3.5 Sonnet: 61% functional correctness, only 10.5% secure. Adding security hints doesn't help. Security needs fundamentally different approaches, not better prompting. (arXiv)
Each link below shares sources, entities, or timing with this story.
I've spent real hours tuning the CLAUDE.md in my own repos. Rewriting architecture notes. Adding conventions. Trimming when it got long. So this one stung. arXiv 2607.27250 ran a two-agent ablation across Claude Code and Codex: 17 real tasks from 3 repositories, 288 gold-test-...
The first systematic study of deceptive UI impact on LLM web agents, accepted at IEEE S&P 2026, tested against real e-commerce, streaming, and news dark patterns. Gemini 2.5 Pro: 65.78% susceptibility. Claude 3.7 Sonnet: 53.79%. GPT-4o: 51.26%. Guardrail models and prompt post...
arXiv 2608.05108 skips the RL-trained attacker models that dominate red teaming and generalize poorly, instead accumulating a strategy library across a sequence of (dataset, target) pairs that transfers to unseen targets with no retraining. AgentDojo: 86.7% ASR against Gemini-...
This is the paper of the week. arXiv 2607.28871 introduces BSG-VA, which replays every validation command an agent runs across three code states: the original buggy code (B), the candidate patch (S), and the gold developer fix (G). If a test passes in all three states, it neve...
A GitHub repo cataloging Claude Code tips doesn't normally warrant a top story. But shanraisshan/claude-code-best-practice at 53.4K stars isn't a tips list anymore. It's the de facto reference for how an entire generation of developers is learning to work with AI coding agents...
Two thirds. Not two thirds of a contrived jailbreak set. Two thirds of realistic malicious issue requests, against the exact three tools most of the people reading this run daily. Ankur Singh, Jinqiu Yang, and Tse-Hsun Chen built IssueTrojanBench across four attack categories...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.