Fetching from the wire…
Research2026-07-31 · source-backed
PAIChecker audits SWE-bench Verified and finds misalignment across five patterns and eleven scenarios, a direct consequence of the construction pipeline pairing a PR with whatever issue its description references, then using the issue as problem statement and the patch as test oracle. Their multi-agent checker hits 92.12% binary accuracy on SWE-Gym and 91.67% on SWE-bench Multilingual across four backbones. Anyone quoting SWE-bench deltas of a couple points should treat 13.6% as a noise floor on the benchmark itself.
Each link below shares sources, entities, or timing with this story.
arXiv 2608.09802, accepted at COLM 2026, audits the benchmark everyone quotes in funding decks and finds a large chunk of the reported headroom is measurement error, not model failure. Their replacement is 170 expert-curated multilingual refactoring instances across Python, Ja...
While the capability stories pile up, here's the counterweight. As SWE-bench Verified scores cluster near saturation on July leaderboards, an enhanced analysis (SWE-Bench+, on the AIware 2026 benchmark track) found 60.83% of commonly resolved issues contain solution leakage ri...
Wu et al. name history reliability as a distinct failure mode: trace entries that stay structurally valid and semantically plausible after they stop being authoritative. On Qwen3-1.7B, polluted history flipped 32.1% of decisions correct under the original trajectory, usually v...
arXiv 2607.27146 attacks from-scratch program synthesis, where agents get only natural-language docs and an execute-only binary as oracle. The pipeline auto-converts open-source command-line programs into source-free training environments and uses GLM-5.2 as teacher for synthe...
Equips code agents with structured memory built from historical commits, distilling intent-to-code mappings with self-refinement via verification feedback. Using DeepSeek-V3.2 as backbone, boosts SWE-bench Verified from 68.4% to 77.8% — new SOTA. Co-evolution with project hist...
RealSWE builds a six-category information taxonomy and four style dimensions, then compares real prompts from SWE-chat against SWE-bench Verified and Pro. Also: 87% of real prompts are casually written, against 94% of benchmark problems written formally. They release 381 multi...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.