Fetching from the wire…
Security2026-08-13 · source-backed
arXiv 2608.11878 replaces the handful of manually implemented injection-testing environments with an Environment Simulator, Attacker Agent, and User Simulator that generate executable stateful environments and discover viable injection points automatically. Injection timing and placement materially change attack success, which the hand-built benchmarks couldn't show. Alignment data from the framework improves security on both ToolHazard-Bench and AgentDojo without hurting benign utility.
Each link below shares sources, entities, or timing with this story.
arXiv 2608.27234 calls the planner exactly once per query to emit a full plan in a declarative DSL, then applies dual-lattice information-flow control over confidentiality and integrity across explicit data flows and control dependencies, storing results as labeled artifacts a...
It treats untrusted data entering agent state as contamination and restricts future capabilities so the state cannot reach deployer-defined forbidden states, using a Skill Impact Graph, steerability signatures and an inline reference monitor (arXiv 2608.30041). Across four Age...
Splitting ASR into covert success, injections leaving no trace in the final response, and overt success, ones a user can spot, follows from the ReAct format where the final response summarizes the most recent action (arXiv 2608.30362). A trace that hands control back to the us...
arXiv 2608.05108 skips the RL-trained attacker models that dominate red teaming and generalize poorly, instead accumulating a strategy library across a sequence of (dataset, target) pairs that transfers to unseen targets with no retraining. AgentDojo: 86.7% ASR against Gemini-...
First defense modeling multi-turn indirect prompt injection as temporal causal takeover. Uses counterfactual re-executions at tool-return boundaries to detect when tool outputs steer agent behavior. Evaluated on AgentDojo across four task suites. Builder-ready pattern for tool...
The Agentic Principal Chain (arXiv 2608.15888) reframes prompt injection as an authorization problem: an injection only matters if the agent holds authority to act on it. Across 3,154 instances spanning InjecAgent, AgentDojo, and ASB, exfiltration hit 0% in all four AgentDojo...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.