A Pre-Registered LLM-Judge Audit Shows Its Own Significant Result Was 79 to 85% Scale Artifact
The standard strongest design for auditing LLM judges, a within-item contrast between two responses differenced again across a manipulated attribute on a bounded rating scale, is not identified on the scale that reports it, because each term is censored by its own share and confounds preference with differential attenuation. In a pre-registered audit of a frozen pedagogy judge sealed before the first of 990 calls, the registered primary endpoint was null at +0.085 points (95% BCa [-0.167, +0.353], p = 0.684). The one nominally significant interaction, +0.378 at p = 0.002, was reproduced 79 to 85% by a construction containing zero differential preference, using only the observed severity shift and the scale floor.
↳ Follow the thread