GoBench puts GPT-6 Astra at 2,568 Elo on 9x9 Go, still 1,800 Elo short of KataGo
Roland Gao published GoBench on 2026-09-15, scoring frontier models on 9x9 Go against a calibrated ladder of KataGo opponents from random to superhuman, with an arXiv release dated 2026-09-17. The leaderboard: GPT-6 Astra Max 2,568, Astra High 2,227, Claude Opus 5 High 2,076, GPT-5.6 Sol Max 1,929, Sol High 1,846, against KataGo's 4,400. Given coding tools and two hours of preparation before evaluation, Codex with Astra reaches 3,560, a 1,000-point jump that is the actually interesting number here. The author claims r=0.83 correlation with ARC-AGI-2; the top r/MachineLearning reply pushes back that Go training data is free to generate, so the benchmark is arbitrary in a way ARC deliberately was not.
↳ Follow the thread