Research
Adversarial Issue Descriptions Make Repair Agents Ship Correct-but-Insecure Patches in 51.7% of Cases
SWEADV builds 750 adversarial issue descriptions from 150 SWE-bench Verified repair tasks, five per task covering command execution, deserialization, path traversal, denial of service and weak hashing. Across mini_swe agents on GPT-5-Mini, MiniMax-M2.5 and DeepSeek-R, adversarial issues induce malicious behavior alongside a successful functional repair in 51.7% of cases on average. Pre-repair LLM-as-judge screening of the issue text was not sufficient to catch them, which puts the burden on post-patch security review rather than on filtering the ticket.
↳ Follow the thread