SourcesAdversarial Intent Latent Variable Stateful Trust for Agentic RAGarXiv·high signalXBlueskyLinkedInCopy linkPOMDP framework for agent security. MMA-RAG Trust Agent achieves 6.5x attack success rate reduction. Most rigorous agent defense paper this week.SourceSource pagearXiv↳ Follow the threadPolicy dependency / Stack layerAgents hallucinate tools that do not exist, a 675B model does it as often as a 7B one, and merging MCP servers adds new failure surfacesarXiv 2609.19425Policy dependency / Stack layerA Cheap Read-Only Verifier Captures Nearly All the False-Pass Benefit of a Full Planning StackarXiv 2609.20474Stack layer / Threat patternInfinite-Parameter LLMs generate feed-forward weights from live data instead of freezing themarXivPolicy dependency / Stack layerEvery Step Passes Its Guardrail and the Workflow Still Violates Policy, and No Step-Scoped Monitor Can Catch ItarXiv 2609.18820Stack layer / Threat patternAdversarial Agents Ran Arbitrary Bash Past Claude Code Auto Mode and Codex Guardian in 79% of TrialsarXiv 2609.19587Stack layer / Threat patternModel-Checking an Agent's Plan Before Any Tool Runs Rejects Unsafe Plans Without Spending a Single Tool CallarXiv 2609.18674Stack layer / Threat patternAgentPProf brings pprof flame graphs to agent trajectories by segmenting on task boundaries instead of call stacksarXiv 2609.20301Stack layer / Threat pattern37,623 provenance-labeled agent PRs: Codex code was reverted half as often as human code, Devin's 31% more, and Claude Code PRs waited 12.6 hours for first reviewarXiv 2609.17598