HorizonMath: Measuring AI Progress Toward Mathematical Discovery with Automatic Verification
arXiv 2603.15617·medium signal
Introduces a benchmark targeting genuinely unsolved mathematical problems — not competition math — with automatic verification of proposed solutions, directly addressing the saturation of existing math benchmarks. Motivated by the gap between benchmark performance and actual research-level discovery as illustrated by the Knuth-Claude interaction. Provides a rigorous framework for evaluating whether AI is approaching genuine mathematical frontier work.