SourcesPyVision-RL Solving Agent Interaction CollapsearXiv·medium signalXBlueskyLinkedInCopy linkAddresses RL-trained agents learning to reduce tool usage. Oversampling-filtering-ranking rollout strategy sustains interaction.SourceSource pagearXiv↳ Follow the threadPolicy dependency / Stack layerHolding Back Ready Agent Turns Instead of Releasing Them Eagerly Cuts P95 Workflow Latency up to 3.50xarXiv 2609.10964Stack layer / Update threadSpecGuard Turns Speculative Decoding's Acceptance Rate Into a Free Runtime Backdoor DetectorarXivStack layer / Threat patternOne in Seven Python Samples That Pass Bandit and Semgrep Still Carries a Runtime-Confirmed ExploitarXivStack layer / ContrastSkill optimization via contextual bandits cut optimization cost 55-58% using only 50 examples per benchmarkarXiv 2609.11682Stack layer / ContrastEvoSafeHarness searches policies and code together to build a per-model safety harness, cutting attack success from 45.6% to 10.0%arXiv (2609.05903)Stack layer / Update thread100 LLM agents running a simulated town economy for 26 weeks froze the money supply: 0.3% of prices ever changed, and memory deletion made no differencearXivStack layer / ContrastEcdysis: fix the harness only for failure patterns that recur across tasks, not for each single failurearXiv 2609.11677Stack layer / Update threadNCP-ArchPreview reaches OLMo-3-7B's final pretraining loss on 51.3% of the tokens by predicting multi-token conceptsarXiv / HuggingFace Daily Papers