ResearchMOSAIC Safe Multi-Step Tool Use Plan-Check-Act FrameworkarXiv·high signalXBlueskyLinkedInCopy linkPost-training framework for safe agent tool use, 50% harmful behavior reduction, 20%+ refusal on injectionSourceSource pagearXiv↳ Follow the threadPolicy dependency / Stack layerRIPPLE: an edit confined to one prompt-policy segment changes downstream behavior, so replay candidate edits after previously accepted ones before persistingarXiv 2609.12127Stack layer / Threat patternUnlearning Methods That Pass TOFU and MUSE Still Leak the Secret on 22-86% of Queries Once the Model Is an AgentarXiv 2609.12808Stack layer / ContrastTwo-gap framework recasts reward hacking and hallucination as symptoms of requirement and model gapsarXivStack layer / Update threadSplitting a CTF Task Into Isolated Sub-Contexts Lets a Local gemma-4 Solve 18.52% of Challenges Standard Agent Loops FailarXiv 2609.12839Policy dependency / ContrastAn audit of a production agent pipeline found a reported p99 latency of 2,147,483,647 ms, the signed 32-bit maximum, from a lifecycle clamparXiv 2609.12017Policy dependency / Follow-up threadMOSAIC Picks a GraphRAG Traversal Policy per Query and Beats the Best Fixed Policy by 9.96 PointsarXiv 2609.11065Stack layer / ContrastOdin Runs All 32 Llama-3-8B Transformer Layers Under FHE in 366 Seconds on One H100, 4.51x Faster Than THORarXiv 2609.12378Stack layer / Follow-up threadBenchmark Radar Ships a Daily-Updated Catalog of 1,283 AI Benchmarks With 12,916 Numeric Score ObservationsarXiv 2609.11115