Threat taxonomy for autonomous pentest agents argues conversational guardrails do not transfer to agents with persistent memory and real-world actions
This security analysis of LLM-powered autonomous penetration testing agents systematically works through representative architectures, characterizes their trust boundaries and attack surfaces, and proposes a lifecycle-aligned threat taxonomy spanning LLM lifecycle attacks, agent-architecture attacks and cross-cutting behavioral attacks. The central claim is that persistent memory, real-world action-taking and long-horizon reasoning create qualitatively different risks from chat-based LLM systems, so existing conversational guardrails are structurally insufficient for offensive-security agents. It is a survey rather than an empirical result, so treat it as a checklist for anyone deploying HexStrike- or ARXON-class tooling rather than as new measurement.
Source
↳ Follow the thread