Policy dependency / Stack layer
Converting GUI Trajectories Into Replayable MCP-Style Calls Instead of Unstructured Memories
arXiv 2609.16635
Stack layer / Threat pattern
37,623 provenance-labeled agent PRs: Codex code was reverted half as often as human code, Devin's 31% more, and Claude Code PRs waited 12.6 hours for first review
arXiv 2609.17598
Stack layer / Contrast
Encoding the Expected Answer as Executable Code Keeps Data-Science Agent Evals Valid on Live Data
arXiv 2609.16487
Stack layer / Contrast
Emergence World ran 10 agents per world for 16 days and found no frontier model contained an injected attack — one acted on poisoned memory 46 hours later
arXiv
Policy dependency / Stack layer
Cognitive Admission Control makes an agent produce a witness certificate before it is allowed to mutate infrastructure
arXiv 2609.16313
Stack layer / Threat pattern
PentestChain keeps a 7B local model off the critical path behind a deterministic exploit map and runs an eleven-tool MCP pentest pipeline at zero paid-API cost
arXiv 2609.18120
Stack layer / Update thread
Two Tool-Level Defenses Drive Prompt Injection and Memory Poisoning to 0% Attack Success in Many Settings
arXiv 2609.16098
Stack layer / Contrast
Giving a coding agent a refreshing visual map of the repo graph gains 2.4 points on SWE-bench Verified while cutting tokens 5.8%
arXiv