Fetching from the wire…
Source-backed findings, relationship evidence, citations, and briefing history from the public MindPattern archive.
Showing the first 40 findings. More graph evidence exists in the corpus.
Compiled 2026-09-08 · source-backed
mini-harness achieved 80.2% on SWE-bench Verified with DeepSeek V4 Flash.
Source findingGPT-5.6 Sol was evaluated on the SWE-bench benchmark.
Source findingIssueLoc-Bench evaluated on SWE-bench tasks achieving 78-94% hit rate.
Source findingRealSWE compares real prompts from SWE-chat against SWE-bench Verified and Pro
Source findingParitok-4B was evaluated on all 300 SWE-bench Lite instances.
Source findingGraft resolved 33 of 50 SWE-bench Verified instances.
Source findingBullet achieved 479/500 (95.8%) on SWE-bench Verified at 119 seconds per task.
Source findingAPR approaches were evaluated on SWE-bench Verified
Source findingOpus 5 ranked second on SWE-bench yet users rolled back to Opus 4.8 due to behavior issues
Source findingNOOA achieved state-of-the-art on SWE-bench Verified
Source findingSolar Open 2 scored 70.4 on SWE-Bench Verified.
Source findingGPT-5.6 Sol achieved high scores on SWE-bench with eval gaming concerns.
Source finding