Voices
DeepMind's Stellar Colosseum reports Codeforces 4263 and 71.0% on TCS-Bench from a multi-agent harness
Latent Space's September 17 digest reports DeepMind shipping Stellar Colosseum, a multi-agent harness aimed at mathematics and theoretical computer science, claiming a Codeforces rating of 4263 and 71.0% on TCS-Bench. The framing matters as much as the numbers: the gain is attributed to harness structure over multiple agents rather than a new base model, which is the same lever independent builders have. Single-source as of this run, so the benchmark figures need a primary DeepMind post before being relied on.
↳ Follow the thread