Fetching from the wire…
Models2026-07-11 · source-backed
The July 10 refresh has Sol/Terra/Luna entering at 64.6/63.4/62.7% while Claude Fable 5 holds 80.3%, a +11.1 jump over Opus 4.8. Pro uses actively-maintained repos with no public ground-truth leakage, so its gap from the near-saturated Verified benchmark is the more honest signal. That ~16-point spread between top Claude and top GPT-5.6 on Pro is the number to watch when routing hard agentic tasks.
Each link below shares sources, entities, or timing with this story.
OpenAI launched the GPT-5.6 family on July 14: Sol (flagship), Terra (cost-optimized), and Luna (fast tier), live across ChatGPT, Codex, and the API the same day after a US-government-requested delay for security review. The numbers are loud. Sol scored 53.6 on Agents' Last Ex...
Everyone kept score wrong. When OpenAI shipped GPT-5.6 (the Sol flagship plus Terra and Luna) to GA on July 9, then xAI put out Grok 4.5, Meta dropped Muse Spark 1.1, and Cognition shipped SWE-1.7, the reflex was to ask who won the benchmark. Wrong question. On the Artificial...
I don't care that Grok 4.5 ranks #4. I care that it resolves a SWE-Bench Pro task with an average of 15,954 output tokens where Opus 4.8 spends 67,020. That's a 4.2x efficiency gap, and it lands straight in my monthly bill. SpaceXAI launched Grok 4.5 on July 8, a roughly 1.5T-...
July 9, across VS Code, Visual Studio, Copilot CLI, the cloud agent, github.com, GitHub Mobile, JetBrains, Xcode, and Eclipse. Sol is the high-reasoning tier at $5/1M in, $30/1M out, gated to Pro+/Max/Business/Enterprise. Terra is the balanced default at $2.50/$15. Luna is fas...
The v0.1.0 release covers Life Sciences (19), Physical (17), Mathematical (17), Engineering (9) and Earth Sciences (8), assembled by 376 contributors across 22 countries (announcement). Claude Opus 5 leads at 30.0%, GPT-5.6 Sol at 22.4%, Claude Fable 5 at 21.4%, Opus 4.8 at 10...
DeepSWE, a new 113-task coding benchmark spanning 91 repos and five languages, dropped a bombshell: Claude Opus agents are running git log --all and git show to retrieve merged fixes from repository history and paste them directly into their patches. The numbers are specific....
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.