Fetching from the wire…
Models2026-08-13 · source-backed
xAI shipped it August 12 with a 500K context, February 2026 cutoff, $2/$6 per million. It scored 61 on the Artificial Analysis Index, tying GPT-5.6 Sol Max, one point behind Fable 5 Max. The number that got 334 points and 381 comments on HN is from Artificial Analysis's teardown: on long-horizon knowledge work it resolves tasks in ~53 turns and ~0.5B input tokens against Opus 5's ~103 turns and ~2.0B, at $0.84 measured cost per task. It loses clearly to Grok 4.5 High on DeepSWE and Terminal-Bench, so the agent-coding framing is uneven. Turn efficiency is the line item that moves if you pay per token.
Each link below shares sources, entities, or timing with this story.
I don't care that Grok 4.5 ranks #4. I care that it resolves a SWE-Bench Pro task with an average of 15,954 output tokens where Opus 4.8 spends 67,020. That's a 4.2x efficiency gap, and it lands straight in my monthly bill. SpaceXAI launched Grok 4.5 on July 8, a roughly 1.5T-...
SpaceXAI released Grok 4.5 on July 8, and for once the vendor hype and the third-party numbers point roughly the same direction. Musk called it "roughly comparable to Opus 4.7, but much faster." Priced at $2 per million input tokens and $6 per million output, that's over 60% b...
Everyone benchmarks per task. Accuracy on SWE-bench, pass rate on Terminal-Bench, a leaderboard row per model. Together AI ran the experiment sideways: fix the budget at $100, point both models at DeepSWE, and count how much work came out the other end. GLM-5.3 finished 17 tas...
The leaderboard says first place. The methodology says you should check your own bill. Qwen3.8 Max now ranks first on Artificial Analysis' agentic index, scoring 86.1 on OSWorld-Verified ahead of GPT-5.6 Sol Max at 83.2 and Fable 5 at 85.0, priced at $2.00/M input and $6.00/M...
Artificial Analysis has Grok 4.6 at 61, one point under Fable 5 Max's 62, leading GDPval-AA v2 at 1753 (vs 1741 and 1728) and AA-Briefcase at 1577 (vs 1574 and 1502), beating Sol on 6 of 9 shared benchmarks. Then Terminal-Bench v3.0: 26% versus 34.6% for Sol and 34.1% for Fabl...
This is the other half of the Fable 5 story, so read them together. While the best coding model in the world is uncallable, an open-weight one quietly posted frontier-adjacent numbers. Per Tom's Hardware, independent benchmarks for the MIT-licensed GLM-5.2 (744B params, 40B ac...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.