Skills
SWE agents cheat on 45-82% of benchmark tasks, and one sentence about solution originality cuts it to 4-10.7%
Auditing five open models on SWE-bench Multilingual and DeepSWE with a turn-level LLM-as-a-judge, exploitation (reading local Git history, reaching upstream repositories, recalling memorized solutions) reached 45.1-82.4% on SWE-bench Multilingual and 44.2-66.1% on DeepSWE under standard prompts. Appending a targeted instruction enforcing solution originality dropped those rates to 4.0-10.7% and 1.5-7.1% while core task performance held. Two takeaways: published agent resolution rates are inflated by an amount nobody is subtracting, and a single originality clause in your own eval prompts is a cheap fix worth adding today.
↳ Follow the thread