Stack layer / Threat pattern
Two-gap framework recasts reward hacking and hallucination as symptoms of requirement and model gaps
arXiv
Stack layer / Threat pattern
A Jailbreak SoK Finds Low Final-Response Attack Success Hides Compromised Planning, Memory, and Tool State
arXiv 2609.12413
Stack layer / Threat pattern
Every Model Tested Edits Already-Optimal Code 100% of the Time, and a Training-Free Guardrail Lifts Abstention to 44.4%
arXiv 2609.14839
Stack layer / Threat pattern
Persistent Memory Poisoning Hits Claude Code at 81.7% Cross-Session Attack Success and OpenClaw at 55.5%
arXiv 2609.13889
Stack layer / Threat pattern
MemRiskBench Scores Long-Horizon Agent Memory Risks Deterministically, With No LLM Judge on the Pass/Fail Path
arXiv 2609.14976
Policy dependency / Threat pattern
An audit of a production agent pipeline found a reported p99 latency of 2,147,483,647 ms, the signed 32-bit maximum, from a lifecycle clamp
arXiv 2609.12017
Policy dependency / Threat pattern
A Fine-Tuned RoBERTa-Large Permission Gate Matches Claude Haiku 4.5 at Deciding What an Agent May Touch
arXiv 2609.15422
Stack layer / Threat pattern
TRAIL pairs a translator agent against a challenger agent and gains 23.1% relative syntax accuracy on C-to-Rust translation
arXiv