ResearchDefensive Refusal Bias Safety Alignment Fails Cyber Defenders 2.72xarXiv·high signalXBlueskyLinkedInCopy linkSafety-tuned LLMs refuse legitimate defensive security tasks at 2.72x rate, system hardening 43.8% refusalSourceSource pagearXiv↳ Follow the threadStack layer / Threat patternThe Post-Training Method, Not the Data, Decides How Refusal Is Computed Inside a ModelarXiv 2609.03887Stack layer / Threat patternOpen-Weight Code Models Fabricate on 60% of Impossible Tasks and Refuse Only 27%arXiv 2609.03267Stack layer / Threat patternAlcaTRAz Defends Jailbreaks With Character-Level Perturbation Rules and No Model Access, Beating Llama Guard on 73.4% of CombinationsarXiv 2609.03693Stack layer / Threat patternUMPeek Recovers Private User Models From a Personalized Agent's Choices, With No Access to Memory or BackendarXiv 2609.03815Stack layer / Threat patternPACE tests whether an assistant will refuse a reasonable-sounding request because of something it has to retrieve about you firstarXivStack layer / Threat patternAligning Latent Moral Representations Instead of Responses Improves Adversarial Robustness Where Behavioral Alignment Made It WorsearXiv 2609.04022Stack layer / Threat patternThe same model scores 0.00 or 0.96 on tool calling depending only on the serving adapterarXivStack layer / Threat patternA Blockchain-Anchored Black Box for Agent Workflows, Explicitly Scoped to Evidence Rather Than PreventionarXiv 2609.04017