Fetching from the wire…
Research2026-08-20 · source-backed
Across five TTS methods and five benchmarks spanning medicine, law, finance, chat and creative writing: candidate generation kept improving with compute in every domain, but reward models correlated with actual quality at roughly ρ=0.12. Only candidate fusion consistently beat single-sample, recovering about 40% of available quality gain. (arXiv 2608.18931) So best-of-N with a reward-model picker is a bad deal on subjective tasks. Synthesize across candidates instead of ranking them.
Each link below shares sources, entities, or timing with this story.
SpecFirst splits the loop in two: a spec agent probes the binary and fuses observations with documentation into a structured specification, then a separate synthesis agent codes against that fixed reference. Test pass rates rose 6.9-21.3% and binary exploration coverage 9.4-18...
RAGAS-style evaluation checks correctness against a frozen snapshot, which means routine document updates and corrections can silently break production without moving a dashboard. This ASE 2026 paper defines 11 mutation operators perturbing at both the pre-chunk index level an...
Writer launched Palmyra X6 on August 13 with a number that should reset how you think about agent COGS: 52% lower average cost, 48% better speed, 10% better quality. The model is a post-training variation of Z.ai's open-source GLM-5.2. A US enterprise SaaS vendor built its fla...
TSDS pairs a convergence probe that halts local reasoning once the intended action stops changing with a perplexity rule that escalates only genuinely uncertain actions to a larger cloud model, jointly tuned by Learn-Then-Test for finite-sample guarantees on both expected epis...
arXiv 2607.12227 (Wang et al., incl. Hajishirzi, Tsvetkov, Dasigi) finds two methodological holes in the self-improving-agent literature: methods are never compared against simpler baselines at matched compute budgets, and final performance gets reported on the same public ben...
SpecPath found 35 of 100 passing implementations broke when only the revision path changed, with aggregate accuracy looking identical across paths. Build your eval set from real multi-turn clarification threads with amendments and reversals. Path sensitivity is invisible to st...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.