Fetching from the wire…
Security2026-08-18 · source-backed
The Agentic Principal Chain (arXiv 2608.15888) reframes prompt injection as an authorization problem: an injection only matters if the agent holds authority to act on it. Across 3,154 instances spanning InjecAgent, AgentDojo, and ASB, exfiltration hit 0% in all four AgentDojo domains and all 544 InjecAgent data-stealing cases were blocked. Intent binding cut destruction from 38.6% to 4.0% and manipulation from 90.5% to 12.1%. The honesty is what sells it: utility drops 8.6 and 13.9 percentage points across 949 task-injection pairs. That's a real trade and they published it instead of burying it.
Each link below shares sources, entities, or timing with this story.
arXiv 2608.11878 replaces the handful of manually implemented injection-testing environments with an Environment Simulator, Attacker Agent, and User Simulator that generate executable stateful environments and discover viable injection points automatically. Injection timing an...
AgentFlow specifies where data may travel through an agent system rather than which individual actions are unsafe, using flow and path rules, task-scoped capabilities, controlled release and stateful taint semantics, enforced by a runtime monitor plus an SMT-based verifier (ar...
First defense modeling multi-turn indirect prompt injection as temporal causal takeover. Uses counterfactual re-executions at tool-return boundaries to detect when tool outputs steer agent behavior. Evaluated on AgentDojo across four task suites. Builder-ready pattern for tool...
arXiv 2608.05108 skips the RL-trained attacker models that dominate red teaming and generalize poorly, instead accumulating a strategy library across a sequence of (dataset, target) pairs that transfers to unseen targets with no retraining. AgentDojo: 86.7% ASR against Gemini-...
A July 21 paper pairs two near-identical agents: an Explore Agent that inspects untrusted input but holds no tools, and a Safe Agent that takes privileged actions using its own context plus length-constrained hints from the explorer (arXiv 2607.19595). Borrowing from residual...
It interposes a lightweight model (GPT-4o, 4.1, or o4-mini) that detects and excises injected instructions from untrusted input before your agent ever sees them, using fuzzy-regex excision of the flagged span (arXiv 2507.15219). On AgentDojo it pushes false positives, false ne...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.