Policy dependency / Stack layer
RIPPLE: an edit confined to one prompt-policy segment changes downstream behavior, so replay candidate edits after previously accepted ones before persisting
arXiv 2609.12127
Stack layer / Contrast
Synthetic Document Finetuning Makes a Model Look Aligned but Fails to Inoculate Against Reward-Hacking Misalignment
arXiv 2609.14998
Stack layer / Update thread
Splitting a CTF Task Into Isolated Sub-Contexts Lets a Local gemma-4 Solve 18.52% of Challenges Standard Agent Loops Fail
arXiv 2609.12839
Stack layer / Contrast
Removing the Tenant ID From an MCP Tool Schema Blocks Cross-Tenant Reads That a Validated Parameter Let Through 26 Times Out of 26
arXiv 2609.14780
Stack layer / Follow-up thread
Ericsson's Multi-Agent Code Reviewer Hits 96% Accuracy, With a Third of Its Correct Findings Rated Must-Fix
arXiv 2609.15877
Policy dependency / Stack layer
Python's Import Statement Is an Execution Boundary: 90% of Initialization-Activated Advisory Vulnerabilities Are High or Critical
arXiv 2609.14791
Stack layer
A Weak Local Model Splits a Harmful Task Into Benign Subproblems and Launders Frontier Capability, Raising a CBRN Rubric Score From 62.3 to 83.1
arXiv 2609.15383
Stack layer
A Jailbreak SoK Finds Low Final-Response Attack Success Hides Compromised Planning, Memory, and Tool State
arXiv 2609.12413