Sources
J-Zero co-evolves a Challenger, Solver and Judge from zero data and gains 8.0 points on unverifiable tasks
Self-evolving models have mostly been demonstrated in verifiable domains where a checker exists. This 2026-08-27 paper extends the loop to unverifiable ones by co-adapting the Judge alongside the adversarial Challenger and Solver pair. The trick is where the Judge's supervision comes from: preference pairs whose ordering is known in advance by construction (the Solver's answer beats the Challenger's, a decomposed-and-recombined answer beats a one-shot answer) rather than from the Judge's own scores, which sidesteps the circularity that usually sinks self-judging. It beats baselines by 4.2 points on verifiable and 8.0 on unverifiable domains and keeps improving across iterations.
↳ Follow the thread