Fetching from the wire…
Public story · 2026-07-30 · high
Four model tiers spanning a 15x price gap failed at the same rate: no model buys its way out of a stale-data problem.
Why now: This entered the same late-July window as a related proposal to check agent tool calls against server-held facts instead of trusting the model's own rationale.
Agents turned stale-but-valid pricing data into costly wrong actions about 60% of the time, per SARC-DQ's July 28 benchmark. On the pricing-replenishment task, that's an agent reordering stock at a superseded price because the record still looks current. Whoever owns that inventory budget eats the cost.
SARC-DQ isn't testing garbage-in problems like malformed JSON or missing fields. It's testing data that's wrong only in freshness, lineage, or provenance, records that pass every schema check regardless. Neither the system's data-quality flags nor the agent's own hedging language caught the stale records better than a coin flip. Both scored an AUC at or below 0.50.
The finding that should reorder some roadmaps: this 60% conversion rate held flat across four model tiers spanning a 15x gap in inference price. The priciest model in the test didn't catch more stale records than the cheapest one did. That's the paper's core claim: evidence integrity is a separate axis from model capability, so a frontier upgrade doesn't fix it.
Its fix is a metadata-aware gate that runs before the agent acts, checking freshness, lineage, and provenance instead of catching errors after the fact. It doesn't say what that gate costs in false positives, or whether the 60% figure holds outside pricing-replenishment specifically.
Each link below shares sources, entities, or timing with this story.
Testing eight models across 192,000 evaluations, researchers found chain-of-thought and direct instructions to ignore the score didn't remove the bias.
A proposed provenance gate cut unauthorized high-risk actions to zero after the attack itself hit a 1.000 success rate in tests.
Five coding harnesses that pass identical tests burn up to ten times more tokens than each other, and extra spend can't recover a fact that's missing.
A 4,181-problem study found confidence-based escalation beats a frozen router by 4.2 points while using 37% fewer tokens.
Within 48 hours, three unrelated sources landed on the same structural problem from three directions.
Evolved harnesses lost to a matched-budget sampling baseline once tested on a benchmark the search never saw.
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.