Fetching from the wire…
Source-backed findings, relationship evidence, citations, and briefing history from the public MindPattern archive.
Showing the first 40 findings. More graph evidence exists in the corpus.
Opus 4.8 was evaluated on CommerceAgentBench at 56/107 tasks.
Source findingClaude Code version 2.1.233 uses Opus 4.8 model
Source findingTodoWrite was removed from Opus 4.8.
Source findingOpus 4.8 scored 84.6 on Terminal-Bench 3.0
Source findingClaude Code on Opus 4.8 achieves 57.2% on SWE-Atlas QnA
Source findingQwen-UI-Agent is competitive with Claude Opus 4.8 on GUI agent benchmarks
Source findingOpus 5 achieves 0% browser prompt-injection success with Auto Mode versus higher rates for Opus 4.8.
Source findingClaude Code with Opus 4.8 scores 78.9% on Terminal-Bench 2.1.
Source findingGrok 4.5 undercuts Opus 4.8 by 60% in cost ($2.49 vs $11.80 per task)
Source findingWillison documents that Opus 4.8 performs worse than older versions with custom edit tools
Source findingWillison attributes Datasette's 2026 code-frequency spike to using Opus 4.8 for development.
Source findingOpus 4.8 is among top performers on SWE-bench Pro.
Source finding