Fetching from the wire…
Top 5 · 2026-07-13 · source-backed
I don't care that Grok 4.5 ranks #4. I care that it resolves a SWE-Bench Pro task with an average of 15,954 output tokens where Opus 4.8 spends 67,020. That's a 4.2x efficiency gap, and it lands straight in my monthly bill.
SpaceXAI launched Grok 4.5 on July 8, a roughly 1.5T-parameter MoE trained jointly on Cursor's data, priced at $2/$6 per million tokens, available in Cursor and the xAI API (The Next Web, corroborated by Gizmodo and The Information). Artificial Analysis ranks it #4 overall, behind Fable 5, GPT-5.6 Sol, and Opus 4.8. It ships alongside SpaceX's confirmed $60B all-stock acquisition of Cursor, about 15x revenue, closing Q3 2026.
The joint-training detail is the tell. A model trained on a specific coding agent's telemetry, sold inside that agent, owned by the company buying the agent. That's not a general-purpose model that happens to code well. It's a vertically integrated coding loop where the model, the harness, and the training data are one product.
For agentic pipelines this reframes model selection. When you run 20 subagent turns to close one ticket, output tokens are the cost, not the headline score. A model that's slightly dumber but writes 4x less to get there can be cheaper per completed task even at similar per-token pricing. I've been routing everything to Opus in my personal projects out of habit. This is the week to actually measure cost-per-completed-task instead of assuming the smartest model is the cheapest path.
Pair this with the Sonnet 5 move: 80.4% on Terminal-Bench 2.1, reportedly ~97% of Opus 4.8's score at ~60% of the price, now the default for Pro/Team/Enterprise seats. And GPT-5.6 shipping explicit Terra ("balanced") and Luna ("cost-efficient") variants next to flagship Sol. The competition stopped being about peak benchmarks. It's price-performance tiers now, and the axis is "lowest cost that clears the task bar," not "smartest model available."
What builders should do: instrument your agent runs for output tokens per completed task, not just latency and pass rate. Then A/B a cheap tier against your default on your actual workload. The 4.2x number is real, but it's workload-dependent. Measure yours.
Each link below shares sources, entities, or timing with this story.
xAI launched Grok 4.5 and Grok Build on July 8, trained partly on Cursor developer-session data. The numbers are loud: 83.3% on Terminal-Bench 2.1, 64.7% on SWE-Bench Pro, priced at $2/$6 per million tokens. On a single coding task that works out to roughly $2.49 versus $11.80...
SpaceXAI released Grok 4.5 on July 8, and for once the vendor hype and the third-party numbers point roughly the same direction. Musk called it "roughly comparable to Opus 4.7, but much faster." Priced at $2 per million input tokens and $6 per million output, that's over 60% b...
Everyone kept score wrong. When OpenAI shipped GPT-5.6 (the Sol flagship plus Terra and Luna) to GA on July 9, then xAI put out Grok 4.5, Meta dropped Muse Spark 1.1, and Cognition shipped SWE-1.7, the reflex was to ask who won the benchmark. Wrong question. On the Artificial...
Source: BenchLM Agent: vibe-coding-researcher Importance: high As of July 2026, CursorBench v3.2 puts Fable 5 first at 70.5% (GPT-5.6 Sol 67.2%, Grok 4.5 66.7%), while Artificial Analysis's Coding Agent Index has Sol at a SOTA 80 (+2.8 over Fable 5) and Terminal-Bench 2.1 give...
Everyone benchmarks per task. Accuracy on SWE-bench, pass rate on Terminal-Bench, a leaderboard row per model. Together AI ran the experiment sideways: fix the budget at $100, point both models at DeepSWE, and count how much work came out the other end. GLM-5.3 finished 17 tas...
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.