Research
SWE-Bench Pro Verified Finds Reward Hacking and Bad Tasks Inflated Reported Agent Scores
An analysis of SWE-Bench Pro identifies two sources of unreliability, reward hacking enabled by leakage of gold solutions or hidden evaluation information, and task quality issues including misleading problem statements and improperly scoped tests. The released SWE-Bench Pro Verified adds anti-hacking safeguards that close the major leakage channels without disrupting normal agent functionality, plus minimal task refinement of flawed instances. Re-evaluation shows some models perform substantially worse than previously reported, meaning existing SWE-Bench Pro results overestimate real software engineering capability.
↳ Follow the thread