Fetching from the wire…
Public story · 2026-07-14 · high
GPT-5.6 Sol Ultra scored highest overall at 91.9%, outside the roughly 35-agent CLI field Morph ranks as of the July 14 standings.
Why now: The July 14 standings give teams a fixed reference instead of re-litigating CLI choice after every model release.
Codex CLI running GPT-5.5 leads the public Terminal-Bench 2.1 leaderboard at 83.4% in the July 14 standings, per Morph.
That's the number worth citing when a team argues over which coding CLI to standardize on. Morph counts roughly 35 actively maintained CLI agents in the same field as of July 14.
GPT-5.6 Sol Ultra scored highest overall at 91.9%, but that sits outside the public CLI-pairing leaderboard Codex CLI leads.
Claude Code paired with Opus 4.8 is the top usable Claude pairing at 78.9%, about 4.5 points behind Codex CLI. Gemini CLI with Gemini 3.1 Pro trails both at 70.7%.
Opencode is the most-starred option among that open-source field, per Morph.
Morph doesn't say what disqualifies a Claude pairing from the top spot, or how excluded configurations scored. So 78.9% is a floor for Claude Code, not the full range of what's possible with it.
A 4.5-point Terminal-Bench gap between Codex CLI and Claude Code is a tiebreaker for whoever's picking cold, not a verdict on whoever's already shipping. If Codex CLI's lead holds through the next model drop for Claude Code, that's the signal worth switching over. Until then, the standings are a reference for new decisions, not grounds to relitigate old ones.
The July 14 standings give teams something fixed to point to instead of restarting the CLI argument after every model release.
Each link below shares sources, entities, or timing with this story.
Terminal-Bench 2.1 results (entries dated June 17) put Codex CLI on GPT-5.5 first at 83.4%, Claude Code on Fable 5 second at 83.1%, and Claude Code on Opus 4.8 at 78.9%. The asterisk matters more than the ranking: Fable 5 and Mythos 5 have been export-suspended since June 12,...
Goose moved to the Linux Foundation, OpenCode relocated to the anomalyco org and now cites 160K+ stars, 900+ contributors, and 7.5M monthly developers (Morph). On capability, Codex CLI with GPT-5.5 still leads Terminal-Bench at 83.4%, while Google retired the consumer Gemini C...
xAI launched Grok Build on May 14. With that, every major AI lab now ships a coding agent that lives in your terminal. The competition isn't "can we build one" anymore. That question is settled. The lineup: Anthropic has Claude Code. OpenAI has Codex CLI. Google has Gemini CLI...
Everyone spent yesterday arguing about benchmark numbers. Tencent quietly published data suggesting the numbers belong to your infrastructure, not the model. The WorkBuddy Bench leaderboard reports every model under two different agent harnesses — CodeBuddy Code and Claude Cod...
Claude Code 2.1.229 shipped a config flag most people will scroll past: CLAUDE_CODE_WORKFLOW_PREFIX_STAGGER_MS. What it does is delay the launch of sibling agents that share a prompt prefix, so the second through Nth agents read the warm cache instead of each writing their own...
43.3% on Frontier-Bench v0.1. Opus 4.8 scored 18.7%. That's not an incremental bump, that's the same benchmark with a different shape of answer. Anthropic released Claude Opus 5 on July 24 at $5/$25 per million input/output tokens, exactly half of Fable 5's $10/$50, while matc...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.