Fetching from the wire…
Models2026-08-14 · source-backed
Artificial Analysis has Grok 4.6 at 61, one point under Fable 5 Max's 62, leading GDPval-AA v2 at 1753 (vs 1741 and 1728) and AA-Briefcase at 1577 (vs 1574 and 1502), beating Sol on 6 of 9 shared benchmarks. Then Terminal-Bench v3.0: 26% versus 34.6% for Sol and 34.1% for Fable 5. It wins professional-work evals and price, loses long-horizon coding agents. The Reddit framing that it "beats Sol" flattens a genuinely split result.
Each link below shares sources, entities, or timing with this story.
SpaceXAI released Grok 4.5 on July 8, and for once the vendor hype and the third-party numbers point roughly the same direction. Musk called it "roughly comparable to Opus 4.7, but much faster." Priced at $2 per million input tokens and $6 per million output, that's over 60% b...
I don't care that Grok 4.5 ranks #4. I care that it resolves a SWE-Bench Pro task with an average of 15,954 output tokens where Opus 4.8 spends 67,020. That's a 4.2x efficiency gap, and it lands straight in my monthly bill. SpaceXAI launched Grok 4.5 on July 8, a roughly 1.5T-...
The leaderboard says first place. The methodology says you should check your own bill. Qwen3.8 Max now ranks first on Artificial Analysis' agentic index, scoring 86.1 on OSWorld-Verified ahead of GPT-5.6 Sol Max at 83.2 and Fable 5 at 85.0, priced at $2.00/M input and $6.00/M...
xAI shipped it August 12 with a 500K context, February 2026 cutoff, $2/$6 per million. It scored 61 on the Artificial Analysis Index, tying GPT-5.6 Sol Max, one point behind Fable 5 Max. The number that got 334 points and 381 comments on HN is from Artificial Analysis's teardo...
Everyone kept score wrong. When OpenAI shipped GPT-5.6 (the Sol flagship plus Terra and Luna) to GA on July 9, then xAI put out Grok 4.5, Meta dropped Muse Spark 1.1, and Cognition shipped SWE-1.7, the reflex was to ask who won the benchmark. Wrong question. On the Artificial...
Terminal-Bench 2.1 results (entries dated June 17) put Codex CLI on GPT-5.5 first at 83.4%, Claude Code on Fable 5 second at 83.1%, and Claude Code on Opus 4.8 at 78.9%. The asterisk matters more than the ranking: Fable 5 and Mythos 5 have been export-suspended since June 12,...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.