Policy dependency / Stack layer
Holding Back Ready Agent Turns Instead of Releasing Them Eagerly Cuts P95 Workflow Latency up to 3.50x
arXiv 2609.10964
Stack layer / Contrast
EvoSafeHarness searches policies and code together to build a per-model safety harness, cutting attack success from 45.6% to 10.0%
arXiv (2609.05903)
Stack layer / Update thread
Market Concentration Barely Changes Model Collapse: Pushing One Model to 90% Share Moves Five-Generation Endpoints a Few Percent
arXiv 2609.11146
Threat pattern / Contrast
BlueSTAR runs tiered autonomous cyber defense on two live enterprise IT/OT ranges
arXiv
Stack layer / Contrast
Skill optimization via contextual bandits cut optimization cost 55-58% using only 50 examples per benchmark
arXiv 2609.11682
Stack layer / Update thread
An Unlearning Audit of 263 Released Checkpoints Finds 47 Move Past Their Own Seed Spread Just by Refitting Batch-Norm Statistics
arXiv 2609.11490
Stack layer / Follow-up thread
Benchmark Radar Ships a Daily-Updated Catalog of 1,283 AI Benchmarks With 12,916 Numeric Score Observations
arXiv 2609.11115
Policy dependency / Follow-up thread
MOSAIC Picks a GraphRAG Traversal Policy per Query and Beats the Best Fixed Policy by 9.96 Points
arXiv 2609.11065