Fetching from the wire…
Source-backed findings, relationship evidence, citations, and briefing history from the public MindPattern archive.
Showing the first 40 findings. More graph evidence exists in the corpus.
SWE-2 scored 92.8% on Terminal-Bench 2.1.
Source findingmini-harness scored 69.44% on Terminal-Bench 2.1.
Source findingK2 Horizon 375B scored 70.2% on Terminal-Bench 2.1, reduced to 66.9% after contamination audit.
Source findingGLM-5.3 scores 88.2 on Terminal-Bench 2.1.
Source findingOrnith 1.5 397B MoE scores 86.1 on Terminal-Bench 2.1.
Source findingDarwinX achieved 83.2% on Terminal-Bench 2.1, up 7.7 points.
Source findingLongHorizon-Harness improved Terminal-Bench 2.1 success to 77.2%
Source findingClaude Code on Sonnet 5 shows the largest degradation at 18.3 points on Terminal-Bench 2.1
Source findingDeepSeek-V4-Flash-0731 scored 82.7 on Terminal Bench 2.1.
Source findingGrok 4.5 achieved 83.3% on Terminal-Bench 2.1
Source findingCodex CLI topped Terminal-Bench 2.1 at 83.4%.
Source findingGLM-5.2 scored 81.0 on Terminal-Bench 2.1.
Source finding