Fetching from the wire…
Public story · 2026-08-16 · high
A hidden-state probe predicts answerability with 0.91 AUROC, but no model scores above 0.292 on the benchmark.
Why now: TRAPSBench posted to arXiv in August 2026, and the paper doesn't explain why models catch textual impossibility about 4 times more easily than missing visual evidence.
TRAPSBench pits 16 vision-language models against 1,404 matched physics video pairs, each one rigged so a single change makes the answer undeterminable. The best model scored 0.292 on a metric built to punish wrong guesses as hard as it rewards right ones, per the arXiv paper. These systems answer confidently even when the video hides the evidence they'd need.
A linear probe run against the models' hidden states predicted whether a question was even answerable, up to 0.91 AUROC, the researchers found. The model's internals knew. Its output didn't say so.
Steering makes the case stronger. Nudging a single "void" direction in one layer of the network causally turns abstention on or off, the paper reports.
Models catch a textually impossible question about 4 times more often than they catch a video missing the same visual evidence. Language gives them an out. Pixels don't.
That looks less like confusion than reluctance. A model that decodes answerability at 0.91 AUROC but scores only 0.292 already has the signal. It just doesn't act on it before answering.
Each link below shares sources, entities, or timing with this story.
Built from 542 quality-controlled factual questions into 6,504 episodes with tool returns of known correctness, across five open-weight 7-9B models (arXiv 2608.26295). Models follow a correct tool 86.0-93.1% of the time and repeat the tool return in 78.4-86.0% of cases where b...
The framing is Gricean: an uncertain cooperative speaker retreats up the specificity hierarchy, trading informativeness for truthfulness. On a T-REx-based benchmark varying entity familiarity and referent specificity, model activations do encode whether a referent falls inside...
HANDBOOK.md is a benchmark for whether standing instructions actually constrain an agent across extended tool-use runs. Not whether the model reads your policy file. Whether it still obeys it forty tool calls deep. 65 tasks pairing expert-written SOPs of 20 to 124 pages across...
Testing five VLMs across two benchmarks and five visual-token budgets, native-resolution table images match text on accuracy and efficiency, but downscaling makes models compensate for lost readability with longer, weaker reasoning traces that cancel the token savings. The exp...
A prespecified randomized audit ran seven models over 3,024 choice sets, three personas, nine paraphrases and nine arms for 40,068 scored responses (arXiv 2608.14399). Reputation dominates, with a 3.9 to 4.7 rating raising choice probability 31.4 points. But demographic parity...
arXiv 2608.13010 scores top-five retrieval candidates against ranks 6–20 of the same query to spot answer-anchor concentration, and separately compares documents to lexically distinct neighbors to catch coordinated density before any query arrives. Deployed jointly, attack suc...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.