TruthInsightBench Withholds the Answer Entirely, and Four Coding Agents Plateau at 58.4-60.3 of 100
arXiv 2609.05079 argues existing AI-scientist benchmarks are configured for reproduction — tasks, data, and rubrics built around a hidden target study whose result recovery is rewarded — which is not discovery. Its 40 blind tasks, drawn from 40 peer-reviewed studies across 10 domains, expose only a neutral scientific objective and frozen data, withholding conclusions, expected values, and analysis paths. A fixed LLM judge scores evidentiary maturity across six dimensions as 29 artifact-grounded items with no per-instance human grading; on one frozen base model, four coding agents cluster at 58.4-60.3 with no statistically reliable pairwise separation, strong on auditability but lacking the discriminating acts that establish a claim.
↳ Follow the thread