Fetching from the wire…
Public story · 2026-03-17 · source-backed
OpenAI shipped the first model family explicitly designed for subagent pipelines. GPT-5.4 mini features a 400K context window, scores 54.4% on SWE-Bench Pro (vs. the flagship's 57.7%), and handles computer use at 72.1% on OSWorld — at $0.75 input / $4.50 output per million tokens, 70% cheaper than base GPT-5.4. Cached input hits $0.075/M.
Nano is the real story for high-volume pipelines. At $0.20/M input and $1.25/M output, it's cheaper than Gemini 3.1 Flash-Lite and outperforms the previous GPT-5 mini at max reasoning effort. Simon Willison benchmarked vision tasks at 0.069 cents per image — 76,000 museum photo descriptions for $52.44 total.
The naming matters: "subagent workloads" is now a first-class product category. Every multi-agent pipeline has the same cost problem — running frontier models at every node burns budget on tasks that don't need frontier capability. Mini handles the thinking nodes; nano handles classification, extraction, and routing. The plan-and-execute pattern (Opus/GPT-5.4 as planner, mini/nano as executors) just got a 90% cost reduction on the executor side.
Mini is available in ChatGPT free tier via Thinking mode. Nano is API-only at launch. For teams already running heterogeneous model routing through tools like ccNexus or claude-launcher, these slot directly into the cheap-executor tier. For teams not yet doing model routing — this is the pricing signal that makes it irrational not to start. Source
Each link below shares sources, entities, or timing with this story.
OpenAI shipped GPT-5.5 on April 23, six weeks after 5.4. The capability jump is real: 82.7% on Terminal-Bench 2.0 vs Claude Opus 4.7's 69.4%. The Pro tier nearly doubles Opus 4.7 on FrontierMath Tier 4 at 39.6% vs 22.9%. It uses 40% fewer tokens on Codex tasks while matching 5...
OpenAI released GPT-5.4 in Standard, Thinking, and Pro variants. Headline capabilities: native computer-use (75.0% on OSWorld-Verified, surpassing human 72.4%), 1M token context, and first-ever "compaction" support for longer agent trajectories. The Tool Search API is the buil...
GPT-5.6 Luna went to $0.20 input / $1.20 output per million tokens on July 30. That's an 80% cut. Terra dropped 20%. Luna's input now undercuts Gemini 3.1 Flash-Lite ($0.25/$1.50) and sits at one-fifth of Claude Haiku 4.5's $1 input. Simon Willison covered the announcement and...
Zhipu's GLM-5.1 ranked third on Code Arena, jumping 90+ points over its predecessor GLM-5 and landing ahead of GPT-5.4 and Gemini 3.1 Pro. Two separate r/LocalLLaMA threads (491 upvotes and 233 upvotes) confirm this isn't just a benchmark curiosity. Practitioners are paying at...
Gemini 3.6 Flash launched July 21 at $1.50/1M in, $7.50/1M out, claiming 17% fewer output tokens than 3.5 Flash, DeepSWE code precision up from 37% to 49%, OSWorld-Verified computer use at 83% (from 78.4%), and a knowledge cutoff finally moved from January 2025 to March 2026....
I don't care that Grok 4.5 ranks #4. I care that it resolves a SWE-Bench Pro task with an average of 15,954 output tokens where Opus 4.8 spends 67,020. That's a 4.2x efficiency gap, and it lands straight in my monthly bill. SpaceXAI launched Grok 4.5 on July 8, a roughly 1.5T-...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.