Fetching from the wire…
Tools2026-08-09 · source-backed
Cursor's August 6 post details a two-stage router: "Compass" assigns each turn a 0–1 complexity score by predicting whether the user will be satisfied, then a taxonomy across domains (backend, frontend, database), tasks (bug fixes, commands, tests), and modifiers (visual changes, product questions) picks the model. The pool assigns Grok as cost-efficient baseline, Sol for planning and codebase comprehension, Opus for execution-heavy and DevOps, Fable for debugging and visual implementation, with Opus 5 recently added. Claims: Auto Intelligence delivers above-Fable satisfaction at 68% lower cost, Auto Balance outperforms Opus 4.8 at 41% lower cost, Compass hits 96% positive signal on highest-confidence tasks versus 71% on lowest. Read it as a blueprint. Predicting satisfaction rather than difficulty is the part I hadn't considered.
Each link below shares sources, entities, or timing with this story.
I don't care that Grok 4.5 ranks #4. I care that it resolves a SWE-Bench Pro task with an average of 15,954 output tokens where Opus 4.8 spends 67,020. That's a 4.2x efficiency gap, and it lands straight in my monthly bill. SpaceXAI launched Grok 4.5 on July 8, a roughly 1.5T-...
Cursor shipped two modes, Auto Intelligence and Auto Balance, routing each request based on production traffic feedback rather than static rules. This is the heterogeneous-routing pattern practitioners have hand-built for a year, expensive model for reasoning, cheap for mechan...
xAI launched Grok 4.5 and Grok Build on July 8, trained partly on Cursor developer-session data. The numbers are loud: 83.3% on Terminal-Bench 2.1, 64.7% on SWE-Bench Pro, priced at $2/$6 per million tokens. On a single coding task that works out to roughly $2.49 versus $11.80...
SpaceXAI released Grok 4.5 on July 8, and for once the vendor hype and the third-party numbers point roughly the same direction. Musk called it "roughly comparable to Opus 4.7, but much faster." Priced at $2 per million input tokens and $6 per million output, that's over 60% b...
xAI shipped it August 12 with a 500K context, February 2026 cutoff, $2/$6 per million. It scored 61 on the Artificial Analysis Index, tying GPT-5.6 Sol Max, one point behind Fable 5 Max. The number that got 334 points and 381 comments on HN is from Artificial Analysis's teardo...
Anthropic released Claude Fable 5.1 on September 1. Claude Code v2.1.257 made it the default Fable model at 17:53 UTC that day, with a 1M-token context window, $10 per million input tokens, $50 per million output, and $0.25 per million on cache reads (claude-code CHANGELOG). B...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.