Fetching from the wire…
Public story · 2026-08-16 · high
The self-check finds most of the improvement it first measured was sampling luck, not a better allocation method.
Why now: The paper posted to arXiv in August 2026 with its full data, code, and pre-registration record attached.
Researchers checked their own test-time budget allocation method for neural combinatorial optimization, and most of the reported gain turned out to be sampling luck. That distinction matters because in-sample evaluation is a common shortcut in this field, and this setup shows it can manufacture a confident result from nothing.
On in-distribution TSP-100 problems, an oracle allocation was computed and scored on the same stored samples. That produced a 2.2 to 2.6% gain across three solvers, POMO, AM, and SymNCO, with confidence intervals excluding zero, per the paper.
Measured on samples the allocation method never touched, the gain disappeared: 0.457%, 0.015%, and -0.512% across the same three solvers. More samples and more test instances didn't shrink the bias. It just sat there.
Not every result collapsed. Under distribution shift, the correction held up: a real 11.5% and 12.0% improvement, confirmed in a pre-registered experiment run separately from the exploratory work.
The paper doesn't say how common this in-sample leak is across other published results. It only shows the leak fully explained the gain in this paper's own in-distribution setup. Data, code, and the pre-registration record are all released, so the audit is checkable rather than asserted.
Each link below shares sources, entities, or timing with this story.
SpecFirst splits the loop in two: a spec agent probes the binary and fuses observations with documentation into a structured specification, then a separate synthesis agent codes against that fixed reference. Test pass rates rose 6.9-21.3% and binary exploration coverage 9.4-18...
arXiv 2607.25886 isolates data-centric research capability by fixing the entire post-training stack so only the agent's data strategy varies. Four frontier agents across six benchmarks. Among searches that continued past the best observed score, 78.26% ended on a lower-scoring...
CausalMix tries to attribute downstream capability gains to specific data sources instead of grid-searching the mix as a hyperparameter. Data-mix selection is one of the highest-leverage and least-transparent training decisions, so a principled attribution method here is pract...
RAGAS-style evaluation checks correctness against a frozen snapshot, which means routine document updates and corrections can silently break production without moving a dashboard. This ASE 2026 paper defines 11 mutation operators perturbing at both the pre-chunk index level an...
Most benchmarks test single-episode solving, and memory benchmarks test fact retention. Neither checks procedural reuse, whether an agent can convert a solved session into a reusable search/debug/verify routine (arXiv). Under a Train/Extract/Test protocol with held-out tasks,...
Open your CLAUDE.md right now. Find the line where you told the agent never to touch production, or never to run rm -rf, or never to commit secrets. That line does nothing. Not "might do nothing under adversarial conditions." Nothing, in the sense that no permission rule, no s...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.