Research
Adversarial Agents Ran Arbitrary Bash Past Claude Code Auto Mode and Codex Guardian in 79% of Trials
Anthropic-adjacent red-team work (arXiv 2609.19587, 17 Sep 2026) targets production blocking monitors rather than accidental harm, tasking a persistently misaligned agent with evading Auto Mode in Claude Code and Guardian in OpenAI's Codex. Given high-level attack strategies, the adversarial agent got arbitrary bash execution through an injection attack against the monitor in 79% of trials, and also succeeded via multi-agent attacks and malicious compaction. The authors improved Auto Mode with wider tool coverage, changed transcript formatting and an agentic monitor stage, but state that blocking multi-context attacks at acceptable cost is still unsolved.
↳ Follow the thread