ResearchDefensive Refusal Bias Safety Alignment Fails Cyber Defenders 2.72xarXiv·high signalXBlueskyLinkedInCopy linkSafety-tuned LLMs refuse legitimate defensive security tasks at 2.72x rate, system hardening 43.8% refusalSourceSource pagearXiv↳ Follow the threadStack layer / Threat patternEmergence World ran 10 agents per world for 16 days and found no frontier model contained an injected attack — one acted on poisoned memory 46 hours laterarXivPolicy dependency / Threat patternA Skill File Can Carry Distilled Memory, Collapsing Two Sets of Retrieval Machinery Into OnearXiv 2609.16669Policy dependency / Threat patternIn a human-operator cyber range, RL-trained defensive agents beat heuristic policies but performance swings with the adversary and the simulated usersarXivPolicy dependency / Threat patternA Fine-Tuned RoBERTa-Large Permission Gate Matches Claude Haiku 4.5 at Deciding What an Agent May ToucharXiv 2609.15422Stack layer / Threat patternTool-Augmented Agents Fabricate Values 45.3% of the Time When a Tool Returns status:ok With an Unusable PayloadarXiv 2609.14758Policy dependency / Stack layerPython's Import Statement Is an Execution Boundary: 90% of Initialization-Activated Advisory Vulnerabilities Are High or CriticalarXiv 2609.14791Stack layer / Threat patternSynthetic Document Finetuning Makes a Model Look Aligned but Fails to Inoculate Against Reward-Hacking MisalignmentarXiv 2609.14998Stack layer / Threat patternA Weak Local Model Splits a Harmful Task Into Benign Subproblems and Launders Frontier Capability, Raising a CBRN Rubric Score From 62.3 to 83.1arXiv 2609.15383