Fetching from the wire…
Public story · 2026-03-06 · source-backed
OpenAI released GPT-5.4 in Standard, Thinking, and Pro variants. Headline capabilities: native computer-use (75.0% on OSWorld-Verified, surpassing human 72.4%), 1M token context, and first-ever "compaction" support for longer agent trajectories. The Tool Search API is the builder-critical feature: models look up tool definitions on demand instead of loading all schemas upfront, reducing token usage by 47%. This directly addresses the context overhead problem — Claude Code hauls 62,600 characters of tool definitions per turn (55% of context), while Tool Search would let models query tools as needed. Benchmarks: 83.0% GDPval (vs Opus 4.6's 78.0%), 57.7% SWE-Bench Pro, 89.3% BrowseComp (Pro variant). GPT-5.2 deprecated in 3 months. (OpenAI | TechCrunch)
Each link below shares sources, entities, or timing with this story.
OpenAI shipped the first model family explicitly designed for subagent pipelines. GPT-5.4 mini features a 400K context window, scores 54.4% on SWE-Bench Pro (vs. the flagship's 57.7%), and handles computer use at 72.1% on OSWorld — at $0.75 input / $4.50 output per million tok...
Martin Alderson's essay "The upcoming AI margin collapse, part 1: GLM 5.2" hit 675 points and 462 comments on Hacker News, and it's the rare HN chart-topper that's actually about spreadsheet math instead of vibes. The argument is simple. Z.ai's GLM 5.2 delivers frontier-adjace...
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's...
New model doesn't mean better model. The r/ClaudeAI community learned this the hard way. The top post on r/ClaudeAI hit 2,757 upvotes with 682 comments calling Opus 4.7 "a serious regression, not an upgrade." Cross-platform sentiment was uniformly negative: 818 upvotes on r/si...
Anthropic launched Claude Sonnet 4.6 claiming performance comparable to Opus 4.5 at $3/$15 per million tokens (vs. Opus's $5/$25). SWE-bench Verified: 79.6% (near Opus 4.6's 80.8%). OSWorld-Verified: 72.5% (tied with Opus 4.6's 72.7%). 1M token context window in beta. Now the...
OpenAI released GPT-5.4 simultaneously across ChatGPT, API, and Codex — the first unified triple release. Built-in computer-use capabilities (build-run-verify-fix loop), 1.05M token context, and 33% fewer false claims vs GPT-5.2. A new experimental "Playwright Interactive" ski...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.