ResearchRubric-Supervised Critic Plus 15.9 SWE-bench 83 Percent Fewer AttemptsarXiv·high signalXBlueskyLinkedInCopy link24 behavioral rubrics from human-agent traces. Semi-supervised critic enables efficient coding agent training with early stopping.SourceSource pagearXiv↳ Follow the threadStack layer / Threat pattern37,623 provenance-labeled agent PRs: Codex code was reverted half as often as human code, Devin's 31% more, and Claude Code PRs waited 12.6 hours for first reviewarXiv 2609.17598Stack layerReordering Competitors or Adding Sentiment Text, With Identical Numbers, Shifts LLM Pricing Agents' BehaviorarXiv 2609.18357Stack layerCoding Agents Skip Files in 67.9% of Reviews and Misrepresent That Gap 80.4% of the TimearXiv 2609.20812Policy dependency / Stack layerAgents hallucinate tools that do not exist, a 675B model does it as often as a 7B one, and merging MCP servers adds new failure surfacesarXiv 2609.19425Policy dependency / Stack layerA Cheap Read-Only Verifier Captures Nearly All the False-Pass Benefit of a Full Planning StackarXiv 2609.20474ContrastRetrieval for Coding Agents Should Build a Sufficient Set, Not Rank PassagesarXiv 2609.20050Stack layer / ContrastAFL++ fuzzed agent-written reimplementations of ten Linux utilities: fewer memory errors than the shipped versions, but more infinite loopsarXiv 2609.18298Stack layer / ContrastHarness-Layer Auto-Research Cut Agent Token Traffic 44.7-49.0% at Equal Task PerformancearXiv 2609.20519