ResearchAgentSentry Temporal Causal Defense Against Indirect Prompt InjectionarXiv·high signalXBlueskyLinkedInCopy linkFirst defense modeling multi-turn IPI as temporal causal takeover. Counterfactual re-execution at tool-return boundaries.SourceSource pagearXiv↳ Follow the threadStack layer / Threat patternA Blockchain-Anchored Black Box for Agent Workflows, Explicitly Scoped to Evidence Rather Than PreventionarXiv 2609.04017Stack layer / Threat patternAlcaTRAz Defends Jailbreaks With Character-Level Perturbation Rules and No Model Access, Beating Llama Guard on 73.4% of CombinationsarXiv 2609.03693Stack layer / ContrastThe Post-Training Method, Not the Data, Decides How Refusal Is Computed Inside a ModelarXiv 2609.03887Stack layer / ContrastAn Upstream Agent's Wrong Claim Captures 54.2% of Answers When It Arrives FirstarXiv 2609.03425Stack layer / ContrastUMPeek Recovers Private User Models From a Personalized Agent's Choices, With No Access to Memory or BackendarXiv 2609.03815Stack layer / ContrastNo model reliably handles factual correction, identity consistency and temporal conflict at once when tools disagree with the userarXiv 2609.03588Stack layer / ContrastPACE tests whether an assistant will refuse a reasonable-sounding request because of something it has to retrieve about you firstarXivStack layer / ContrastSpeculative Macro Commit Pre-Executes Multi-Action Chains and Cuts Agent Wall Time 44.9% on AppWorldarXiv 2609.03236