No Automatic Evaluator Satisfies All Correctness Assumptions, and Equal Aggregate Scores Hide Opposite Behaviors
arXiv 2609.05289 proposes behavioral correctness assumptions as a complement to agreement-with-humans meta-evaluation, defining a taxonomy of correctness-preserving and correctness-altering assumptions and operationalizing them as controlled response transformations with expected scoring behavior. Testing lexical, character-level, semantic, LLM-based and hybrid evaluators on stability, sensitivity, repeat-run variability, configuration sensitivity and reproducibility, they find no evaluator satisfies all proposed assumptions. Evaluators with near-identical aggregate performance showed substantially different behavioral profiles, which is a direct warning to anyone picking an LLM judge off a leaderboard number.
↳ Follow the thread