SourcesVESPO — Stable Off-Policy LLM Training 152 HF UpvotesarXiv·high signalXBlueskyLinkedInCopy linkVariational sequence-level soft policy optimization. Stable training at 64x staleness ratio. Top paper on HuggingFace Feb 23.SourceSource pagearXiv↳ Follow the threadPolicy dependency / Stack layerRIPPLE: an edit confined to one prompt-policy segment changes downstream behavior, so replay candidate edits after previously accepted ones before persistingarXiv 2609.12127Policy dependency / Stack layerPre-registered ablation shows removing an LLM verifier stage from an offensive-security agent shifts median reported findings from 0 to 2 per runarXivPolicy dependency / Threat patternA Fine-Tuned RoBERTa-Large Permission Gate Matches Claude Haiku 4.5 at Deciding What an Agent May ToucharXiv 2609.15422Policy dependency / Stack layerTRAIL pairs a translator agent against a challenger agent and gains 23.1% relative syntax accuracy on C-to-Rust translationarXivPolicy dependency / Stack layerReconstructing Facts From Feed-Forward Residual Vectors Answers Single-Fact Questions in Two-Million-Token ContextsarXiv 2609.12686Policy dependency / Stack layerAtria Dawn's authors studied 769 task records from their own build, and participants called a third of the AI-assisted tasks infeasible without AIarXiv / HuggingFace Daily PapersPolicy dependency / Stack layerZGCM-1 is a fully open 7B dense model whose training cluster was operated by agent swarmsarXiv / HuggingFace Daily PapersPolicy dependency / Stack layerHazardAuditor runs Claude Code, Codex, Hermes and OpenClaw in one harness and normalizes their events to train a guard modelarXiv / HuggingFace Daily Papers