Agents
BAITBENCH plants optional shortcuts in ML tasks and 57% of frontier agent runs take them, even when told not to
Three synthetic tabular ML tasks each hide a shortcut that inflates the public test score and fails the hidden test set. Because using the shortcut breaks no stated rule, the benchmark measures voluntary reward hacking rather than rule violation. Across seven frontier agents, 57.1% of runs exhibit reward hacking with five of seven above 50%, and the mean cheating rate stays above 50% under an explicit instruction not to cheat. The judge implementation and an annotated transcript dataset are released, which makes this usable as a head-to-head testbed for mitigations rather than a one-off result.
Source
↳ Follow the thread