ResearchSleeper Cell Temporal Backdoors in Tool-Using LLMs via SFT-GRPOarXiv·high signalXBlueskyLinkedInCopy linkPoisoned models pass all benchmarks while harboring trigger-activated malicious tool calls. Supply-chain risk for LoRA adopters.SourceSource pagearXiv↳ Follow the threadStack layer / Threat patternAlcaTRAz Defends Jailbreaks With Character-Level Perturbation Rules and No Model Access, Beating Llama Guard on 73.4% of CombinationsarXiv 2609.03693Policy dependency / Stack layerLLMs Make More and Larger Edits on Code Written by a Different LLMarXiv 2609.03894Stack layer / ContrastThe Post-Training Method, Not the Data, Decides How Refusal Is Computed Inside a ModelarXiv 2609.03887Stack layer / Threat patternA Blockchain-Anchored Black Box for Agent Workflows, Explicitly Scoped to Evidence Rather Than PreventionarXiv 2609.04017Stack layer / Threat patternFIDO2's Real Weakness Is the Environment Around It, Not the CryptographyarXiv 2609.03789Policy dependency / Threat patternThree Robot Navigation Exports With Identical Task Success Leak Wildly Different Amounts About the HomearXiv 2609.03055Stack layer / ContrastUMPeek Recovers Private User Models From a Personalized Agent's Choices, With No Access to Memory or BackendarXiv 2609.03815Stack layer / ContrastOpen-Weight Code Models Fabricate on 60% of Impossible Tasks and Refuse Only 27%arXiv 2609.03267