Fetching from the wire…
Source-backed findings, relationship evidence, citations, and briefing history from the public MindPattern archive.
Showing the first 40 findings. More graph evidence exists in the corpus.
GLM-5.3-Flash scored within half a point of Claude Opus 4.8 on coding benchmarks at one-tenth the price.
Source findingClaude Opus 4.8 scored 10.5% on Terminal-Bench-Science.
Source findingClaude Opus 4.8 scored 10.5% on Terminal-Bench-Science.
Source findingOrnith-1.5 397B achieves 86.0 SWE-bench Verified compared to Claude Opus 4.8's 85.8
Source findingClaude Opus 4.8 and GPT-5.6 Sol Ultra were tested on research tasks.
Source findingOrnith-1.5 was benchmarked against Claude Opus 4.8
Source findingClaude Opus 5 scored 24% on SlopCodeBench versus Opus 4.8's 6%.
Source findingClaude Opus 4.8 was developed by Anthropic.
Source findingStateAct improves Claude Opus 4.8 from 20.6% to 26.9% on computer use tasks.
Source findingClaude Opus 4.8 scores 0.886 on SWE-bench Verified.
Source findingClaude Opus 4.8 scored 56 on the Intelligence Index.
Source findingAnthropic released Claude Opus 4.8 which scores 55.7% on the Intelligence Index
Source finding