Fetching from the wire…
Public story · 2026-08-31 · high
Testing 30 models found the confidence a model states out loud barely lines up with whether its answer is right, and fine-tuning widens the gap.
Why now: The paper posted August 31, 2026.
Researchers tested whether a model's own stated confidence, the number you get when you ask how sure it is, actually predicts whether it's right. Across 30 models from three families, they compared that self-report against two harder benchmarks: logits-based confidence on 8 classification tasks, and semantic entropy on 2 generation tasks.
The self-reported number ranks answers roughly right. More often than not, correct answers get higher confidence than wrong ones. But that ranking is weak on average, and it only firms up on easier questions and with stronger base models. On harder tasks or weaker models, the self-report and actual correctness barely track each other.
Instruction-tuned models make this worse. They report higher confidence overall, and sometimes rank answers slightly better, but the gap between what they claim and what they deliver widens. Calibration gets worse even as the ranking gets marginally better. That cuts against the assumption that tuning makes a model more trustworthy about its own limits.
Prompt design doesn't fix it either. Changing how you ask for confidence shifts the distribution of numbers a model reports, not how well those numbers line up with reality. Cueing the model with attitude, framing the question as if confidence matters, inflates the reported number without making it more accurate.
The paper's own recommendation is narrow. Use self-reported confidence to sort candidate answers against each other, never as a cutoff for auto-approving or rejecting one. An app with a slider or badge showing a model's confidence, letting users treat 90 percent as a green light, is running on a mismatch this paper measured directly. The gap doesn't close by tuning the model further. It's built into what self-report is.
Each link below shares sources, entities, or timing with this story.
A fleet evaluation across 46 endpoints from six vendors found a recognition-enforcement gap: source-format features are linearly decodable from activations and models verbally identify forged authority when asked, but some configurations still emit the conflicting tool call. A...
On a verifiable protein-function characterization task routed across tools, model choice swamped federation topology, RL-versus-LLM harness, and prompt expertise: Opus at roughly 92 to 94%, o4-mini at 40 to 50%. Federation across institutional boundaries cost almost nothing (a...
arXiv 2607.24174 (July 27) generated adversarial log entries from real attack traces and got multiple state-of-the-art LLMs to classify traces containing clear indicators of compromise as benign. The defensive gift: the natural-language explanations emitted alongside the class...
arXiv 2608.09643 trains one linear probe per model on paired vulnerable/fixed Python functions, then tests on real disclosed CVEs whose weakness type the probe never saw. Across five open-weight models, the probe ranks the vulnerable function above its fix 61–67% of the time,...
Add a PromptFoo eval suite that runs every prompt change against fixed test cases across multiple models and fails the build on regression (Lakera). Prompt edits become reviewable diffs instead of blind tweaking. Exactly like unit tests, because that's what they should be.
Participants with AI assistance got 9% of answers right. The baseline group, no assistance, got 27%. Three times worse. And their self-reported confidence went from 30% to 76%. The research comes from University of Milano-Bicocca, École Normale Supérieure, and Sapienza, and th...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.