Fetching from the wire…
Public story · 2026-08-05 · high
Apps pinned to the alias inherited a new model with no opt-in, and response time nearly halved.
Why now: The swap landed on August 5, the same day xAI published the announcement citing Artificial Analysis' benchmark comparison against GPT-Realtime-2.1 and Gemini 3.1 Flash.
xAI swapped the model behind grok-voice-latest on August 5, replacing Think Fast 1.0 with 2.0 for anyone pinned to that alias, per xAI. Pinning to that alias means always running xAI's latest voice model, not a fixed version.
Developers building voice features on that alias got new latency and cost behavior with no upgrade prompt. First-audio response time dropped from 1.25 seconds to roughly 0.70 seconds. Anyone who tuned prompts or timeouts against the old model is now running against different numbers.
Think Fast 2.0 scores 82.9% on Artificial Analysis' speech-to-speech benchmark, up from 75.7% for 1.0, per Artificial Analysis figures cited in xAI's announcement. That puts it ahead of GPT-Realtime-2.1 at 79.1% and Gemini 3.1 Flash at 69.5%.
Median reasoning-token use falls to 0.4x the prior baseline, and audio still costs $0.08 a minute. The model reasons in parallel with speech, so tool calls typically fire before it finishes its first sentence, per xAI.
Each link below shares sources, entities, or timing with this story.
OpenRouter is running an exclusive extra 50% on top of Google's own 50% introductory cut, landing Gemini 3.7 Flash at $0.375 per million input and $1.875 per million output. Google's standing Vertex price through end of 2026 is $0.75 and $3.75, doubling to $1.50 and $7.50 on J...
The changelog offers Z.ai's open-weights coding model with a 1M-token context free via Blackbox AI on AI Gateway, default for new eve agents, switchable for existing ones with eve set --model zai/glm-5.2. Excludes Fast mode and the glm-5.2-fast variant. Separately, Gemini 3.7...
The leaderboard says first place. The methodology says you should check your own bill. Qwen3.8 Max now ranks first on Artificial Analysis' agentic index, scoring 86.1 on OSWorld-Verified ahead of GPT-5.6 Sol Max at 83.2 and Fable 5 at 85.0, priced at $2.00/M input and $6.00/M...
Willison's August 13 release adds Gemini 3.7 Flash, 3.6 Flash and 3.5 Flash-Lite plus gemini-embedding-2 and -001, rebuilding on LLM 0.32's structured message and streaming APIs so reasoning, tool calls and results emit as typed stream events while preserving Gemini thought si...
Google shipped it August 13: FrontierCode 1.1 at 43.6% (from 34.4%), DeepSWE v1.1 at 65.3% (from 49.0%), WebDev Arena Elo 1588 (from 1538), AutomationBench enterprise workflow completion at 30.4% (from 17.0%). Introductory pricing is $0.75/M input and $3.75/M output through De...
GPT-5.6 Luna went to $0.20 input / $1.20 output per million tokens on July 30. That's an 80% cut. Terra dropped 20%. Luna's input now undercuts Gemini 3.1 Flash-Lite ($0.25/$1.50) and sits at one-fifth of Claude Haiku 4.5's $1 input. Simon Willison covered the announcement and...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.