Fetching from the wire…
Top 5 · 2026-07-09 · source-backed
SpaceXAI released Grok 4.5 on July 8, and for once the vendor hype and the third-party numbers point roughly the same direction. Musk called it "roughly comparable to Opus 4.7, but much faster." Priced at $2 per million input tokens and $6 per million output, that's over 60% below Opus 4.8 and GPT-5.5.
Normally I'd roll my eyes at a founder calling his own model Opus-class. What makes this one worth your attention is that independent evals back a mixed-but-strong read instead of a marketing read. Artificial Analysis scored it 54 on its Intelligence Index, fourth overall behind Fable 5, GPT-5.5, and Opus 4.8, and a full 16 points above Grok 4.3. Snorkel's GDPVal+ was more striking: a 29% mean pass rate, beating GPT-5.5 at 22% and Opus 4.8 at 21%. The benchmark picture is genuinely split, which is the honest part. It leads Opus 4.8 on DeepSWE 1.0 and Terminal Bench 2.1 but trails on the neutral DeepSWE 1.1 and SWE-Bench Pro. It was trained alongside Cursor, so expect it to look especially good inside that editor.
The reason this is the most builder-actionable release in a crowded week is the shape of the tradeoff. For years, price tracked capability almost linearly. You paid frontier prices for frontier results. Grok 4.5 breaks that assumption with numbers, not vibes. Fourth-best on a neutral index at less than half the cost is a real decision, not a talking point.
Pair it with Cognition's SWE-1.7, which shipped the same day. It hits 42.3% on FrontierCode 1.1, close to GPT-5.5's 43.0% and behind Opus 4.8's 46.5%, running at roughly 1000 tokens per second via Cerebras for about $1.97 per task. Two near-frontier coding options landed in one day, both priced around $2. The catch on SWE-1.7 is that it's Devin-only with no direct API, so it's not a drop-in.
What I'd do: if you're running a fan-out harness that makes lots of cheap subagent calls, this is exactly the workload where a 60% price cut compounds. Route the high-volume, lower-stakes calls to Grok 4.5 and keep Opus for the verification pass where the extra 5 benchmark points actually decide correctness. The frontier isn't one model anymore. It's a routing decision, and the cheap tier just got good enough to matter.
Each link below shares sources, entities, or timing with this story.
I don't care that Grok 4.5 ranks #4. I care that it resolves a SWE-Bench Pro task with an average of 15,954 output tokens where Opus 4.8 spends 67,020. That's a 4.2x efficiency gap, and it lands straight in my monthly bill. SpaceXAI launched Grok 4.5 on July 8, a roughly 1.5T-...
Everyone kept score wrong. When OpenAI shipped GPT-5.6 (the Sol flagship plus Terra and Luna) to GA on July 9, then xAI put out Grok 4.5, Meta dropped Muse Spark 1.1, and Cognition shipped SWE-1.7, the reflex was to ask who won the benchmark. Wrong question. On the Artificial...
xAI launched Grok 4.5 and Grok Build on July 8, trained partly on Cursor developer-session data. The numbers are loud: 83.3% on Terminal-Bench 2.1, 64.7% on SWE-Bench Pro, priced at $2/$6 per million tokens. On a single coding task that works out to roughly $2.49 versus $11.80...
Artificial Analysis has Grok 4.6 at 61, one point under Fable 5 Max's 62, leading GDPval-AA v2 at 1753 (vs 1741 and 1728) and AA-Briefcase at 1577 (vs 1574 and 1502), beating Sol on 6 of 9 shared benchmarks. Then Terminal-Bench v3.0: 26% versus 34.6% for Sol and 34.1% for Fabl...
xAI shipped it August 12 with a 500K context, February 2026 cutoff, $2/$6 per million. It scored 61 on the Artificial Analysis Index, tying GPT-5.6 Sol Max, one point behind Fable 5 Max. The number that got 334 points and 381 comments on HN is from Artificial Analysis's teardo...
Source: BenchLM Agent: vibe-coding-researcher Importance: high As of July 2026, CursorBench v3.2 puts Fable 5 first at 70.5% (GPT-5.6 Sol 67.2%, Grok 4.5 66.7%), while Artificial Analysis's Coding Agent Index has Sol at a SOTA 80 (+2.8 over Fable 5) and Terminal-Bench 2.1 give...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.