SourcesSAGE — Self-Aware Guided Efficient Reasoning Reduces Inference CostarXiv·high signalXBlueskyLinkedInCopy linkByteDance. Reasoning models know when to stop thinking but sampling obscures it. SAGE selects concise correct paths via confidence.SourceSource pagearXiv↳ Follow the threadPolicy dependency / Stack layerHolding Back Ready Agent Turns Instead of Releasing Them Eagerly Cuts P95 Workflow Latency up to 3.50xarXiv 2609.10964Stack layer / ContrastSkill optimization via contextual bandits cut optimization cost 55-58% using only 50 examples per benchmarkarXiv 2609.11682Policy dependency / Follow-up threadMr.LHDR benchmark: the best deep-research agent gets 43.1% of final answers right but only 34.3% with a correct intermediate chainarXivStack layerLooped Flows Train Recurrent Reasoning With Local Denoising and Hit 58.8% on ARC-AGI-1arXivStack layerBenchShield raises reward-hacking detection in agent benchmarks from 23-94% to 77-100% full-chain recall, using a lifecycle model of reward eventsarXivStack layerThe Power Flexibility Index Measures How Much Throughput an LLM Training Job Loses When You Cut Its PowerarXiv 2609.11542Stack layerVikingRAG Matches State-of-the-Art Structured-Document RAG Using 5.1-32.5% of the TokensarXivPolicy dependencyTencent study: agents cross authorization boundaries in 55-62% of runs when a control constraint is missing and an unsafe action is executable, and 87% when compaction drops the constraintarXiv