Fetching from the wire…
Public story · 2026-08-17 · high
The review warns that gains from a better model rarely carry into end-to-end results, so swapping models to test a harness fix skews the outcome.
Why now: The synthesis is part of the August 17, 2026 briefing.
A 314-page multivocal review argues that coding-agent reliability comes down to harness design, not model capability, per arXiv 2608.13867. For teams debugging flaky agents, that shifts where to look. The review counts 206 reliability records, 193 of them gated practices, aimed at execution state, retrieval, and memory rather than model swaps.
Those 206 records come from 164 academic papers, 100 practitioner accounts, 29 benchmark records, and 17 case studies. Of the 193 gated practices, 56 got developed in real depth. The review also packages 13 research leads, 5 reusable agent skills with evidence maps, and runnable evaluation protocols. Its own advice is to mine the material, not read it cover to cover.
Its sharpest warning is about benchmarking. Gains at one layer, a smarter model or a better retrieval step, don't reliably show up in end-to-end results. The paper's stance is direct. Never test a harness change by swapping the model underneath it.
Each link below shares sources, entities, or timing with this story.
Three frontier models shipped in a single week this month, and teams with a standing eval harness had a routing decision in hours. Anthropic's own agent-eval guidance says 20-50 tasks drawn from your real usage and real failures is enough to detect issues (DeepEval). DeepEval...
Raffi Khatchadourian's replay benchmark measures behavioral instability through three channels that need no access to hidden reasoning text: tool-call trajectories, evidence contacts, decision concentration (arXiv 2607.20491). Across 8,127 replay episodes over 10 models and 3...
A June 25 paper (arXiv:2606.25899) argues manipulation capability varies sharply by task rather than being one measurable scalar. (arXiv) That complicates any safety eval trying to score persuasion as a global number. For anyone deploying agents, the practical implication is t...
Microsoft researchers curated 101 tasks from two production repositories for AL, the DSL behind Dynamics 365 Business Central, adapting SWE-bench's method to an ecosystem with scarce public training data. In bug-fixing, differences between frontier models exceeded differences...
Li, Huo, and Johnson show that one-way message flow between agents produces neither mimicry nor solo behavior but an entirely novel dynamical state, at identical temperature settings. It's conceptual rather than quantitative, but the implication for orchestrator-worker fan-out...
$3,054 against $38,370. Same benchmark, better score. Praxist (arXiv 2608.25955, submitted August 26) replaces per-attempt agent memory with a typed evidence graph of findings, plus lane-structured frontiers and agendas, so later attempts inherit validated mechanisms rather th...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.