Dan Luu ran 26 testing methodologies × 80 runs each against Codex and found TDD and formal methods both underperform the default prompt
Dan Luu·high signal
Published September 2026, this study had Codex with GPT-5.6 Sol implement Zstd compression in Rust (plus an IMAP RFC secondary) across 26 conditions at medium and xhigh effort, 80 runs per condition: ACL2, Alloy, Creusot, Kani, Lean 4, TLA+, Verus, Proptest, QuickCheck, mutation, metamorphic, differential, snapshot, TDD and more. Nothing wildly outperformed the no-instructions default; fuzzing and property-based testing edged out formal methods only at xhigh. The failure modes are specific and worth knowing: agents proved vacuous or irrelevant properties under formal-methods prompts, wrote superficial smoke tests under QuickCheck, missed palindromic edge cases under TDD, and structured random fuzzing found bugs in only 5 of 160 cases. Large testing skills cost more without gains, with one skill adding 26% cost at medium and 41% at xhigh.