Fetching from the wire…
Top 5 · 2026-04-25 · source-backed
OpenAI shipped GPT-5.5 on April 23, six weeks after 5.4. The capability jump is real: 82.7% on Terminal-Bench 2.0 vs Claude Opus 4.7's 69.4%. The Pro tier nearly doubles Opus 4.7 on FrontierMath Tier 4 at 39.6% vs 22.9%. It uses 40% fewer tokens on Codex tasks while matching 5.4's latency. Native desktop navigation baked in. Clicking buttons, typing text, multi-step workflows out of the box. OpenAI
The price doubles too. Standard API hits $5/$30 per million input/output tokens. Pro: $30/$180. OpenAI is betting the capability gap justifies the premium. Meanwhile, DeepSeek V4-Pro launched the same week at $1.74/$3.48 per million tokens. That's a 6-8x cost difference while V4-Pro posts 80.6% on SWE-bench, roughly matching GPT-5.5's comparable score. VentureBeat
Here's what caught me off guard. OpenAI's official prompting guide for 5.5 tells you to throw away everything you learned about prompt engineering for GPT-4 and GPT-5. Define target outcomes and constraints. Let the model pick the path. New text.verbosity controls at the API level steer response length programmatically. Simon Willison tested it with his pelicans benchmark and found the default lagged behind 5.4, but the xhigh reasoning effort level improved things dramatically at the cost of way more tokens.
The market is bifurcating. Frontier labs raise prices banking on capability. Open-weight models close the gap at commodity rates. For builders running production agent pipelines, model routing isn't optional anymore. Use GPT-5.5 or Opus 4.7 for hard reasoning. Route everything else to V4-Flash at $0.14/M input tokens. If you're not routing by task difficulty, you're burning money.
Each link below shares sources, entities, or timing with this story.
OpenAI shipped the first model family explicitly designed for subagent pipelines. GPT-5.4 mini features a 400K context window, scores 54.4% on SWE-Bench Pro (vs. the flagship's 57.7%), and handles computer use at 72.1% on OSWorld — at $0.75 input / $4.50 output per million tok...
DeepSeek released V4 Preview on April 24 with two open-weight variants: V4-Pro (1.6T total parameters, 49B activated via MoE) and V4-Flash (284B parameters, 13B activated). Both support 1M-token context windows. Both are Apache 2.0 licensed. Both are live right now on Hugging...
DeepSeek released V4 on April 24 and the numbers demand attention. V4-Pro is 1.6 trillion parameters total with 49 billion active, MIT-licensed, native 1M-token context. It scores 80.6% on SWE-bench Verified, putting it within 0.2 points of Claude Opus 4.6. On Terminal-Bench 2...
Martin Alderson's essay "The upcoming AI margin collapse, part 1: GLM 5.2" hit 675 points and 462 comments on Hacker News, and it's the rare HN chart-topper that's actually about spreadsheet math instead of vibes. The argument is simple. Z.ai's GLM 5.2 delivers frontier-adjace...
GPT-5.6 Luna went to $0.20 input / $1.20 output per million tokens on July 30. That's an 80% cut. Terra dropped 20%. Luna's input now undercuts Gemini 3.1 Flash-Lite ($0.25/$1.50) and sits at one-fifth of Claude Haiku 4.5's $1 input. Simon Willison covered the announcement and...
The open-weight race just changed constraint. Moonshot AI suspended all new consumer subscriptions on July 20, roughly 48 hours after Kimi K3 launched, because request volume pushed its compute cluster to capacity. Remaining GPUs are reserved for existing paid subscribers. Tec...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.