BenchLM's Eval Registry Reaches 296 Benchmarks and Stops Zero-Filling Missing Scores
BenchLM's continuously-updated state-of-benchmarks registry, last refreshed July 14, 2026, now tracks 296 benchmark definitions across 10 categories. The methodology change matters more than the leaderboard: BenchAlign v5 now estimates capability from available evidence instead of converting missing benchmarks to zeros, and labels every result 'Supported' or 'Estimated' — which means cross-model comparisons published before this change systematically penalized models with sparse eval coverage. Current listing puts Claude Fable 5 at 83 and GPT-5.6 Sol at 81.5 on 'Supported' evidence, with MiniMax M3 leading open-weight at 68.8; the top entry, Claude Mythos 5 at 85.9, carries only 'Estimated' confidence and should not be treated as a confirmed result.
Source
↳ Follow the thread