Research
Breaking Agent Backbones (ICLR 2026): Threat Snapshot Abstraction for LLM Agent Security
Published at ICLR 2026, this paper addresses the core challenge of AI agent security modeling: non-deterministic LLM outputs make fixed execution flows unmappable from inputs, rendering traditional threat modeling inapplicable. The authors introduce 'threat snapshot abstraction' that captures full single-call context alongside attacker objective and method, enabling systematic attack vector categorization. Provides a practitioner-applicable framework for reasoning about where in an agent pipeline an attacker can influence behavior.
↳ Follow the thread