Fetching from the wire…
Agents2026-08-28 · source-backed
arXiv 2608.27234 calls the planner exactly once per query to emit a full plan in a declarative DSL, then applies dual-lattice information-flow control over confidentiality and integrity across explicit data flows and control dependencies, storing results as labeled artifacts and exposing only semantic metadata to later planning. Attack success drops to 0% on AgentDojo and 0.2% on their new multi-query extension. Plan-once-then-enforce keeps showing up as the design that actually works, at the cost of losing mid-run adaptability.
Each link below shares sources, entities, or timing with this story.
arXiv 2608.11878 replaces the handful of manually implemented injection-testing environments with an Environment Simulator, Attacker Agent, and User Simulator that generate executable stateful environments and discover viable injection points automatically. Injection timing an...
arXiv 2608.05108 skips the RL-trained attacker models that dominate red teaming and generalize poorly, instead accumulating a strategy library across a sequence of (dataset, target) pairs that transfers to unseen targets with no retraining. AgentDojo: 86.7% ASR against Gemini-...
A July 21 paper pairs two near-identical agents: an Explore Agent that inspects untrusted input but holds no tools, and a Safe Agent that takes privileged actions using its own context plus length-constrained hints from the explorer (arXiv 2607.19595). Borrowing from residual...
First defense modeling multi-turn indirect prompt injection as temporal causal takeover. Uses counterfactual re-executions at tool-return boundaries to detect when tool outputs steer agent behavior. Evaluated on AgentDojo across four task suites. Builder-ready pattern for tool...
StepGuard (arXiv 2608.24777) is an open-weight guard model auditing individual tool calls pre-execution rather than scoring completed trajectories, which is where most guardrails sit. It trains on paired safe and unsafe trajectories that share identical context and diverge onl...
AgentFlow specifies where data may travel through an agent system rather than which individual actions are unsafe, using flow and path rules, task-scoped capabilities, controlled release and stateful taint semantics, enforced by a runtime monitor plus an SMT-based verifier (ar...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.