Policy dependency / Stack layer
SpecGuard Turns Speculative Decoding's Acceptance Rate Into a Free Runtime Backdoor Detector
arXiv
Policy dependency / Stack layer
A2ABreak extracts a 37-state machine from the A2A spec and finds 11 protocol-level flaws that a fully compliant attacker can exploit
arXiv
Policy dependency / Stack layer
EvoSafeHarness searches policies and code together to build a per-model safety harness, cutting attack success from 45.6% to 10.0%
arXiv (2609.05903)
Policy dependency / Stack layer
Holding Back Ready Agent Turns Instead of Releasing Them Eagerly Cuts P95 Workflow Latency up to 3.50x
arXiv 2609.10964
Policy dependency / Stack layer
100 LLM agents running a simulated town economy for 26 weeks froze the money supply: 0.3% of prices ever changed, and memory deletion made no difference
arXiv
Policy dependency / Stack layer
Φ-Bench tests whether LLMs can engineer their own serving and training stack, from kernels to end-to-end optimization
arXiv
Policy dependency / Stack layer
Skill optimization via contextual bandits cut optimization cost 55-58% using only 50 examples per benchmark
arXiv 2609.11682
Policy dependency / Stack layer
NCP-ArchPreview reaches OLMo-3-7B's final pretraining loss on 51.3% of the tokens by predicting multi-token concepts
arXiv / HuggingFace Daily Papers