Fetching from the wire…
Source-backed findings, relationship evidence, citations, and briefing history from the public MindPattern archive.
Showing the first 40 findings. More graph evidence exists in the corpus.
Claude Opus 4.5 was evaluated against GraphWalker for model-based test path generation.
Source findingClaude Opus 4.5 achieved 98.5% success rate on ALFWorld with EvoHarness-RL.
Source findingClaude Opus 4.5 scores 0.624 on CTI-REALM
Source findingCUDA Agent beats Claude Opus 4.5 by 40% on GPU kernels
Source findingCUDA Agent outperforms Claude Opus 4.5 by 40% on hardest kernel generation tasks.
Source findingClaude Opus 4.5 achieves 80.9% on SWE-Bench Verified coding tasks.
Source findingClaude Opus 4.5 fails 75% of complex workplace tasks on APEX-Agents benchmark.
Source findingClaude Opus 4.5 achieves 37.4% task success on EnterpriseOps-Gym.
Source findingClaude Opus 4.5 scored 80.9% on SWE-Bench Verified.
Source findingClaude Opus 4.5 scores 45.9% on SWE-bench Pro under SEAL standardized scaffolding.
Source findingOpus 4.7 scores lower than Opus 4.5 on SimpleBench
Source findingClaude Opus 4.5 agents closed more deals in Anthropic's Project Deal marketplace experiment.
Source finding