Fetching from the wire…
Public story · 2026-07-14 · high
Grok Build hit 83.3% on Terminal-Bench 2.1 using about a quarter of the output tokens Opus 4.8 needs for the same coding task.
Why now: Grok 4.5 launched July 8, making this the first direct price comparison against Claude's Opus 4.8.
xAI priced a Grok Build coding task at $2.49 on July 8, against $11.80 for Claude Code on Opus 4.8, per TechTimes. That price gap is what budget-conscious teams will care about most. Grok Build runs at roughly a fifth of Claude Code's cost per task, on scores close to frontier level.
Grok Build hit 83.3% on Terminal-Bench 2.1 and 64.7% on SWE-Bench Pro. It got there on about 15,954 output tokens, a quarter of the roughly 67,020 tokens Opus 4.8 used on the same task.
The catch, flagged separately from the benchmark scores: hallucination rates are up. A cheap task that comes back as a confident wrong patch you merge anyway isn't cheap at all.
There's a second bill. The Grok Build CLI was caught silently uploading entire repos, Git history included, to an xAI cloud bucket. The $2.49 sticker price doesn't include what it costs when your codebase leaves the building without you choosing to send it.
The move isn't switching wholesale. Run Grok on work where a hallucination is cheap to catch, refactors with solid test coverage, mechanical migrations, throwaway prototypes.
Keep the high-stakes, hard-to-verify work on the model you trust. My bet: the repo-upload story costs Grok Build more adoption than the hallucination rate does.
Trust in data handling is harder to win back than trust in a single patch. Watch whether xAI changes the CLI's default upload behavior, or teams route around it instead.
Each link below shares sources, entities, or timing with this story.
I don't care that Grok 4.5 ranks #4. I care that it resolves a SWE-Bench Pro task with an average of 15,954 output tokens where Opus 4.8 spends 67,020. That's a 4.2x efficiency gap, and it lands straight in my monthly bill. SpaceXAI launched Grok 4.5 on July 8, a roughly 1.5T-...
SpaceXAI released Grok 4.5 on July 8, and for once the vendor hype and the third-party numbers point roughly the same direction. Musk called it "roughly comparable to Opus 4.7, but much faster." Priced at $2 per million input tokens and $6 per million output, that's over 60% b...
43.3% on Frontier-Bench v0.1. Opus 4.8 scored 18.7%. That's not an incremental bump, that's the same benchmark with a different shape of answer. Anthropic released Claude Opus 5 on July 24 at $5/$25 per million input/output tokens, exactly half of Fable 5's $10/$50, while matc...
Every coding agent leaderboard number you've seen was produced in conditions your security team would reject on sight. Researchers from Accomplish AI and NYU open-sourced Boundary-Bench on August 5. The setup: 12 frontier agent harnesses, roughly 10,000 runs, 89 Terminal-Bench...
Everyone spent yesterday arguing about benchmark numbers. Tencent quietly published data suggesting the numbers belong to your infrastructure, not the model. The WorkBuddy Bench leaderboard reports every model under two different agent harnesses — CodeBuddy Code and Claude Cod...
OpenAI shipped GPT-5.5 on April 23, six weeks after 5.4. The capability jump is real: 82.7% on Terminal-Bench 2.0 vs Claude Opus 4.7's 69.4%. The Pro tier nearly doubles Opus 4.7 on FrontierMath Tier 4 at 39.6% vs 22.9%. It uses 40% fewer tokens on Codex tasks while matching 5...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.