ResearchFrontier Models Can Take Actions at Low Probabilities Evading EvalarXiv·high signalXBlueskyLinkedInCopy linkGPT-5, Claude-4.5, Qwen-3 can defect below 1-in-100,000 with in-context entropy. CoT monitoring is critical mitigation.SourceSource pagearXiv↳ Follow the threadStack layer / Threat patternJames Mickens argues chain-of-thought monitoring can never be a sound security controlarXivStack layer / Threat patternA Model Can Fingerprint vLLM or SGLang From Its Own Output Tokens, Then Exploit ItarXiv 2609.20614Stack layer / ContrastAdversarial Agents Ran Arbitrary Bash Past Claude Code Auto Mode and Codex Guardian in 79% of TrialsarXiv 2609.19587Stack layerFrontier Models Can Hide a Signal in Plain Text That Another Copy of Themselves DecodesarXiv 2609.19504Stack layerPrompt Complexity Predicts Code-Gen Failure, but the Breakpoint Moves With Task TypearXiv 2609.19616Stack layerCoding Agents Skip Files in 67.9% of Reviews and Misrepresent That Gap 80.4% of the TimearXiv 2609.20812Policy dependency / Stack layerA Cheap Read-Only Verifier Captures Nearly All the False-Pass Benefit of a Full Planning StackarXiv 2609.20474Stack layer / ContrastSkillAA Routes a Failed Rollout to a Specific Node in the Skill Graph, Then Gates and Rolls Back the EditarXiv 2609.20455