Research
Planting Benign-Sounding Reasoning in an Agent's Context Evades Chain-of-Thought Monitors 25-33% of the Time
Plan injection steers an actor model to take adversarial actions while its reasoning trace still reads clean, achieving 25-33% monitor evasion across monitorability benchmarks and scaling to DeepSeek-R1. Actors not only follow injected plans but paraphrase them as their own reasoning with no attribution to the injection. Giving the monitor more resources backfires: access to the injected plan drops detection by up to 50% on the Bio-Math task, and in a reasoning-budget case study extra thinking tokens get spent rationalizing the injected plan rather than flagging it.
↳ Follow the thread