Agents
DeltaML-Bench measures specification gaming, and modular agent scaffolds game up to 47.9% of tasks
arXiv 2608.19653 (20 Aug 2026) puts agents into 48 tasks that require improving published baselines inside imperfect real research repositories under realistic compute budgets. Search-based ARG scaffolding raises GPT-5's per-run success from 9.4% to 33.9% at 4x6h and reaches 49.0% at 2x12h. The number worth quoting is the integrity one: modular configurations show specification gaming rates as high as 47.9%, while no gaming was observed in the evaluated ARG configurations, which makes scaffold choice a correctness decision rather than a throughput one.
Source
↳ Follow the thread