Policy dependency / Stack layer
RIPPLE: an edit confined to one prompt-policy segment changes downstream behavior, so replay candidate edits after previously accepted ones before persisting
arXiv 2609.12127
Stack layer / Threat pattern
Unlearning Methods That Pass TOFU and MUSE Still Leak the Secret on 22-86% of Queries Once the Model Is an Agent
arXiv 2609.12808
Policy dependency / Stack layer
CodeBLEU Scored 91% for Both RAG Strategies While One of Them Hallucinated APIs 56.4% of the Time
arXiv 2609.12464
Stack layer / Contrast
Reflexion-Style Verbal Memory Sometimes Lowers Success Versus Plain Retry, and Replay Experiments Show Why
arXiv 2609.12404
Stack layer / Threat pattern
787,562 Function Pairs Show AI Code Is Half the Size of Human Code With Different Defect Classes, Not Fewer
arXiv 2609.12708
Stack layer / Update thread
AIM Adds Index-Level Access Control to Agent Memory and Ships MUMBench for Multi-User Memory Operations
arXiv 2609.12320
Stack layer / Contrast
Bash alone beat typed tool catalogs by 21.8-24.5 points on TheAgentCompany while using up to 72% fewer tokens
arXiv 2609.11999
Stack layer / Update thread
Splitting a CTF Task Into Isolated Sub-Contexts Lets a Local gemma-4 Solve 18.52% of Challenges Standard Agent Loops Fail
arXiv 2609.12839