Research
On Post-2025 Models, Verbalized Confidence Now Beats Logprobs as the LLM-as-Judge Soft Score
arXiv 2609.10996 tests up to 18 LLMs on SummEval, AggreFact and HelpSteer2. It finds the standard advice to prefer log-probabilities over stated confidence no longer holds on top proprietary models released after 2025, and calls this a 'compatibility shift'. Adding an overconfidence advisory and self-debate improved calibration and score spread with little accuracy cost on new models, while pre-2025 models paid a penalty. Accuracy-only reporting hides the shift, so judge pipelines tuned on older models may be using the wrong signal.
Source
↳ Follow the thread