Fetching from the wire…
Public story · 2026-07-27 · high
His fix: keep the AI harness, the model, and your memory layer separate so no single provider's exit sinks the business.
Why now: Nadella's comments landed July 27, the same day Vercel added regional pinning to AI Gateway.
Satya Nadella said on July 27 that companies routing everything through one AI lab may not survive, per TechCrunch. His argument: hand a lab your most sensitive business context, and it can turn around and compete with you using it. His fix is an orchestration layer, keeping the harness separate from the model and memory separate from both.
OpenAI already showed what that competition looks like. Per The Register, OpenAI began selling Presence straight to enterprises on July 22, with no self-serve tier and no public pricing.
Deployment is scoped by OpenAI's own forward-deployed engineers. Every CCaaS and conversational-AI vendor built on OpenAI's models now competes with its own supplier, on services margins instead of software margins.
The same day, the gateway pattern got easier to deploy under real compliance constraints. Per Vercel's changelog, a new inferenceRegion parameter on AI Gateway pins inference to US or EU data centers. That closes a GDPR gap that had been an open request since May.
There's a credential-reselling risk nobody's pricing in yet. Matt Lenhard investigated the relay market for Vectoral; Simon Willison covered it July 26. Operators pool LLM credentials and undercut official API pricing, funded by abused free trials and stolen cards.
Willison's conclusion: there's now an ecosystem hunting for unprotected endpoints, and his fix is strict per-key spending caps scoped to time windows. Route through a gateway, pin your region, cap your keys. That's four hours of work against a supplier that might become your competitor.
Each link below shares sources, entities, or timing with this story.
GPT-5.6 Luna went to $0.20 input / $1.20 output per million tokens on July 30. That's an 80% cut. Terra dropped 20%. Luna's input now undercuts Gemini 3.1 Flash-Lite ($0.25/$1.50) and sits at one-fifth of Claude Haiku 4.5's $1 input. Simon Willison covered the announcement and...
Simon Willison found it in the OpenRouter price list. 1.6T MoE, ~49B active, 1M context, up to 384K output, at $0.435/M input (cache miss) and $0.87/M output. The agentic-coding deltas versus the preview are the story: DeepSWE 12.8 → 62.7, CyberGym 52.7 → 83.3, Terminal Bench...
Thibault Sottiaux at OpenAI published an investigation into "a handful of reports where GPT-5.6 unexpectedly deleted files," finding it happens most commonly when full access mode is enabled in Codex. Simon Willison relayed it. A frontier lab publishing a first-party post-mort...
Martin Alderson's essay "The upcoming AI margin collapse, part 1: GLM 5.2" hit 675 points and 462 comments on Hacker News, and it's the rare HN chart-topper that's actually about spreadsheet math instead of vibes. The argument is simple. Z.ai's GLM 5.2 delivers frontier-adjace...
Simon Willison doesn't hand out superlatives. So when he writes that Z.ai's GLM-5.2 is "probably the most powerful text-only open weights LLM," that's worth stopping for. His June 17 evaluation walks through a 753B-parameter Mixture-of-Experts model with 40B active params, a 1...
1. Set Up Cursor Automations (intermediate) — Event-driven agents from PagerDuty/GitHub/Slack triggers with isolated sandboxes. Cursor Blog 2. Apply Context Engineering to Cut Agent Costs 60-80% (advanced) — Hierarchical token budgets, dynamic tool filtering (max 15), automati...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.