Cheap Models Wrote Spec-Conformant Java That Was Correct 12.9% of the Time
arXiv 2609.18052 (submitted 16 Sep 2026) had Gemini Flash 3, GPT-5.4 mini and Claude Haiku 4.5 solve 992 algorithmic problems as Java Spring Boot service methods against a mandated signature and DTO spec, with iteration forbidden and hardcoded answers banned, producing 7,593 methods and 7,936 measured requests. Structural conformance approached ceiling, but 38.4% of methods do not compute the value they return and only 12.9% of returned answers were correct; methods that genuinely computed answered least often and were correct 19.3% of the time. The inverse relationship between response reliability and correctness is the finding that matters: a cheap model that always answers is the one least likely to be right.
↳ Follow the thread