Fetching from the wire…
Skills2026-07-30 · source-backed
SpecFirst splits the loop in two: a spec agent probes the binary and fuses observations with documentation into a structured specification, then a separate synthesis agent codes against that fixed reference. Test pass rates rose 6.9-21.3% and binary exploration coverage 9.4-18.5% across four models, all statistically significant, on a benchmark where frontier models solve under 1%. The gain comes from resolving documentation ambiguity before any code exists.
Each link below shares sources, entities, or timing with this story.
UOJ-Bench uses real competitive-programming submissions to test generation, error-finding, and repair. In single-attempt evaluation, top models fail to identify errors in over 50% of incorrect submissions (arXiv). Test-time scaling pushes success above 90%, but models also fla...
Across five TTS methods and five benchmarks spanning medicine, law, finance, chat and creative writing: candidate generation kept improving with compute in every domain, but reward models correlated with actual quality at roughly ρ=0.12. Only candidate *fusion* consistently be...
SpecPath found 35 of 100 passing implementations broke when only the revision path changed, with aggregate accuracy looking identical across paths. Build your eval set from real multi-turn clarification threads with amendments and reversals. Path sensitivity is invisible to st...
Writer launched Palmyra X6 on August 13 with a number that should reset how you think about agent COGS: 52% lower average cost, 48% better speed, 10% better quality. The model is a post-training variation of Z.ai's open-source GLM-5.2. A US enterprise SaaS vendor built its fla...
This paper audits its own measurement, which almost nobody does. On in-distribution TSP-100, an oracle budget allocation computed and evaluated on the same stored samples reports a 2.2 to 2.6% gain with intervals excluding zero across POMO, AM, and SymNCO. Measured out of samp...
CodeGrep found BM25 at 0.375 precision actively degrades agent performance, Jina at 0.445 is neutral, and only 0.677 helps. Below roughly 0.45 precision you're paying tokens to make the agent worse. Test on your own repo before you assume retrieval is free upside.
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.