Fetching from the wire…
Research2026-08-28 · source-backed
The v0.1.0 release covers Life Sciences (19), Physical (17), Mathematical (17), Engineering (9) and Earth Sciences (8), assembled by 376 contributors across 22 countries (announcement). Claude Opus 5 leads at 30.0%, GPT-5.6 Sol at 22.4%, Claude Fable 5 at 21.4%, Opus 4.8 at 10.5%, GPT-5.6 Terra at 8.6%, GLM 5.3 at 8.1%, Kimi K3 and Grok 4.6 at 7.1%, GPT-5.6 Luna at 3.3%. The near-3x gap between Opus 5 and Opus 4.8 is far wider than the same pair shows on saturated coding benchmarks, which is what an unsaturated benchmark looks like.
Each link below shares sources, entities, or timing with this story.
I don't care that Grok 4.5 ranks #4. I care that it resolves a SWE-Bench Pro task with an average of 15,954 output tokens where Opus 4.8 spends 67,020. That's a 4.2x efficiency gap, and it lands straight in my monthly bill. SpaceXAI launched Grok 4.5 on July 8, a roughly 1.5T-...
This is the other half of the Fable 5 story, so read them together. While the best coding model in the world is uncallable, an open-weight one quietly posted frontier-adjacent numbers. Per Tom's Hardware, independent benchmarks for the MIT-licensed GLM-5.2 (744B params, 40B ac...
Weeks after launch, Z.ai's open-weight GLM-5.2 now accounts for roughly 75% of all Z.ai model traffic on OpenRouter, with at least one provider serving it past 125 tokens per second (GIGAZINE, citing OpenRouter). The numbers behind the surge: an Artificial Analysis Intelligenc...
Terminal-Bench 2.1 results (entries dated June 17) put Codex CLI on GPT-5.5 first at 83.4%, Claude Code on Fable 5 second at 83.1%, and Claude Code on Opus 4.8 at 78.9%. The asterisk matters more than the ranking: Fable 5 and Mythos 5 have been export-suspended since June 12,...
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's...
The July 10 refresh has Sol/Terra/Luna entering at 64.6/63.4/62.7% while Claude Fable 5 holds 80.3%, a +11.1 jump over Opus 4.8. Pro uses actively-maintained repos with no public ground-truth leakage, so its gap from the near-saturated Verified benchmark is the more honest sig...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.