Fetching from the wire…
Models2026-08-29 · source-backed
18 search API products across 8 providers, scored as the equal-weighted mean of DeepSearchQA F1, BrowseComp exact-answer accuracy and AA-Omniscience accuracy. Perplexity medium leads at 80 for $91.39 per 1,000 tasks at 28.6s each; Parallel advanced and Brave LLM-context tie at 75; the model-only baseline scores 33 at $2.92 per 1,000 tasks. That spread is the first public number putting a price on retrieval quality, and 33 is higher than I'd have guessed for a model answering from memory alone. (Artificial Analysis)
Each link below shares sources, entities, or timing with this story.
On Latent Space July 28, OpenAI core product engineering lead Akshay Nathan said Codex and ChatGPT Work combined reached 10 million users within two weeks of the July 9 launch, with monthly actives up more than 10x since January 2026. The number that should reframe your produc...
A practitioner running a private trivia set found 3.8 failing questions 3.6 answered reliably, at every quantization and sampling setting tried, then checked Artificial Analysis' Omniscience evaluation and found the same regression in offline no-tool knowledge accuracy. The to...
The Hrazdan facility opened August 8, scaling to 300 megawatts and 70,000+ NVIDIA Rubin and Blackwell GPUs by end of 2027, built on NVIDIA DSX (40% more GPUs on the same footprint) with Dell PowerEdge, Schneider Electric power and Vertiv cooling. NVIDIA intends to invest, foll...
GPT-5.6 Luna went to $0.20 input / $1.20 output per million tokens on July 30. That's an 80% cut. Terra dropped 20%. Luna's input now undercuts Gemini 3.1 Flash-Lite ($0.25/$1.50) and sits at one-fifth of Claude Haiku 4.5's $1 input. Simon Willison covered the announcement and...
Three points behind GLM-5.3 at 60, tying GPT-5.6 Terra and Muse Spark 1.2, at $0.09 per task against $0.68 for GLM-5.3 max (Latent Space). It burned 149M output tokens to run the index, of which 134M were reasoning tokens, more than Kimi K3 at 133M or Qwen3.8 2.4T A95B at 136M...
Artificial Analysis has the newly released Alibaba model at the top of its cohort: 27B reasoning model, 256K context, image input, Apache 2.0 permitting commercial use. The caveat in the eval data matters more than the headline: it emitted 160M output tokens across the benchma...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.