Research
FormalTCS: Best Model Scores 11.5 on Autoformalizing Research Claims Versus 28.6 Pass@8 When Handed the Formal Statement
The benchmark builds 175 expert-validated instances from STOC, FOCS, SODA, and COLT papers accepted in 2025-2026, preserving each paper's definitions, assumptions, and proof dependencies with verified Lean formalizations. Translating a natural-language claim into a formal theorem statement is the sharpest bottleneck at 11.5 for the best model, against 28.6 Pass@8 for proving statements humans already formalized. An automated research pipeline built on the benchmark generated 64 new claims of which only 6 survived expert evaluation and proof verification, pointing at research taste as a second barrier beyond formalization.
↳ Follow the thread