Fetching from the wire…
Public story · 2026-08-27 · high
Testing eight models across 192,000 evaluations, researchers found chain-of-thought and direct instructions to ignore the score didn't remove the bias.
Why now: As of August 27, the paper's numbers are the clearest evidence yet that judge-in-the-loop refinement needs isolation between passes.
A prior score biases the re-grade, according to a study spanning eight LLM judges and 192,000 attempted evaluations. Kapetanovic et al.'s paper on prior-score anchoring fed models a prior score, a revision index, or an attempt count as context before asking for a fresh rating. The new score dragged toward the old one in seven of eight models tested. The 95% bootstrap intervals stayed below zero, with an effect size (Cohen's d) reaching 0.71.
The damage shows up in real grading. On categorical industry data checked against human ground truth, judges shown their prior score missed 48% of the corrections they should have caught. They also flipped 10.18% of already-correct judgments to the wrong label.
Chain-of-thought reasoning didn't fix it. Telling the model outright to ignore the metadata didn't fix it either. The anchor held through both.
That matters for multi-pass grading systems, the kind that re-check a prior verdict or score a draft, revise it, then score it again. If the judge model sees what it said last time, it defends the earlier number instead of testing it fresh. It misses the corrections it should make almost half the time.
Each link below shares sources, entities, or timing with this story.
The errors trace back to how the benchmark pairs pull requests with GitHub issues, not just to model quality.
Four model tiers spanning a 15x price gap failed at the same rate: no model buys its way out of a stale-data problem.
An 8,135-trial study finds skill files mostly lock in a procedure, and a 100-item skill pool nearly kills retrieval accuracy.
A 2,420-trial test found a 50:50 mix of relevant and irrelevant items beat an all-relevant AI prompt, per an arXiv paper on agent token costs.
A new arXiv paper finds pretraining gains flip into losses past an optimal context length, as models learn to lean on text instead of memory.
Five coding harnesses that pass identical tests burn up to ten times more tokens than each other, and extra spend can't recover a fact that's missing.
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.