Fetching from the wire…
Source-backed findings, relationship evidence, citations, and briefing history from the public MindPattern archive.
Showing the first 40 findings. More graph evidence exists in the corpus.
Spark-X2.5-4B scored 44.4 on SWE-Bench Pro.
Source findingSWE-Prime fine-tuning on 10% of trajectories achieved 12.2% improvement on SWE-Bench Pro.
Source findingQwen3.8-27B achieved 61.7% on SWE-Bench Pro.
Source findingGemini 3.5 Flash-Lite beats Gemini 3 Flash on SWE-Bench Pro at 54.2% vs 49.6%.
Source findingGrok 4.5 achieved 64.7% on SWE-Bench Pro
Source findingOpus 4.8 is among top performers on SWE-bench Pro.
Source findingClaude Fable 5 leads SWE-bench Pro among current models.
Source findingGPT-5 achieved 14.9% on SWE-bench Pro.
Source findingClaude Opus 4.1 achieved 17.8% on SWE-bench Pro.
Source findingGLM-5.1 achieves state-of-the-art 58.4% on SWE-Bench Pro.
Source findingClaude Opus 4.5 scores 45.9% on SWE-bench Pro under SEAL standardized scaffolding.
Source findingGPT-5.3-Codex scores 57% on SWE-bench Pro.
Source finding