Fetching from the wire…
Research2026-08-12 · source-backed
Built on a post-transformer BDH architecture that reasons recurrently in latent space, it hit 29.5% pass@2 on public ARC-AGI-1 at a computed cost roughly 11x cheaper per task than GPT-5.6 Luna Low, even after OpenAI's 80% price cut on July 30 (Pathway). Be precise about what this claims: 29.5% is well below frontier accuracy. The claim is intelligence-per-dollar, and the frontier being pushed is the cost-accuracy Pareto curve, not the accuracy ceiling. Still interesting. 150 million parameters.
Each link below shares sources, entities, or timing with this story.
Four frontier models. Five sealed engineering problems. The result everybody will quote is that Claude Fable 5 won. The result that should actually change how you work is buried three-quarters down the page. JuliaHub published an evaluation on July 30 running four frontier mod...
OpenAI shipped the first model family explicitly designed for subagent pipelines. GPT-5.4 mini features a 400K context window, scores 54.4% on SWE-Bench Pro (vs. the flagship's 57.7%), and handles computer use at 72.1% on OSWorld — at $0.75 input / $4.50 output per million tok...
OpenAI released GPT-5.4 simultaneously across ChatGPT, API, and Codex — the first unified triple release. Built-in computer-use capabilities (build-run-verify-fix loop), 1.05M token context, and 33% fewer false claims vs GPT-5.2. A new experimental "Playwright Interactive" ski...
poetiq.ai — Detailed technical breakdown of how a $40K-hardware startup achieved 54% on ARC-AGI-2 (vs. Google's 45% at nearly 3x the cost). The key innovation: "learned test-time reasoning" — an iterative refinement meta-system where solutions are generated, receive structured...
The number that reframes everything isn't ten. It's two thousand. OpenAI published "Ten advances in mathematics and theoretical computer science" on August 1, claiming an internal version of Astra produced new results on ten problems that had seen no progress on the main resul...
GPT-5.6 Luna went to $0.20 input / $1.20 output per million tokens on July 30. That's an 80% cut. Terra dropped 20%. Luna's input now undercuts Gemini 3.1 Flash-Lite ($0.25/$1.50) and sits at one-fifth of Claude Haiku 4.5's $1 input. Simon Willison covered the announcement and...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.