ResearchMOSAIC safe multi-step tool useWweb·medium signalXBlueskyLinkedInCopy linkPost-training framework: plan-check-act/refuse loop with preference RL. 50% harmful behavior reduction, 20%+ refusal on injection attacks, preserves benign performance. Tested Qwen2.5-7B, Qwen3-4B, Phi-4. arXiv 2603.03205↳ Follow the threadPolicy dependency / Stack layerRIPPLE: an edit confined to one prompt-policy segment changes downstream behavior, so replay candidate edits after previously accepted ones before persistingarXiv 2609.12127Stack layer / ContrastTwo-gap framework recasts reward hacking and hallucination as symptoms of requirement and model gapsarXivPolicy dependency / Stack layerHazardAuditor runs Claude Code, Codex, Hermes and OpenClaw in one harness and normalizes their events to train a guard modelarXiv / HuggingFace Daily PapersPolicy dependency / Stack layerSeqMoE Reaches 80.22% of Full-Load MoE Performance With Only 45% of Experts Resident in Device MemoryarXiv 2609.12978Policy dependency / Stack layerHuang's actual position on pacing: keep labs in control, use many evaluators not one, and stop publishing p(doom) numbersThe TribunePolicy dependency / Stack layerA Bun packaging quirk left a vulnerable undici in Cline after the CVE was supposedly remediatedGitHubStack layer / ContrastTRAIL pairs a translator agent against a challenger agent and gains 23.1% relative syntax accuracy on C-to-Rust translationarXivStack layer / Update threadn8n 2.40.0 preserves empty-text Anthropic thinking blocks across tool calls and enforces execution timeouts on stuck queue jobsGitHub