Fetching from the wire…
Public story · 2026-07-14 · high
Chinese open-weight models went from 4.5% to 30-46% of US enterprise API traffic in a year, per CNBC.
Why now: Goldman reportedly issued the recommendation to Wall Street clients on July 13, turning a traffic-share trend into an explicit vendor call from a compliance-bound bank.
Goldman Sachs reportedly told Wall Street clients on July 13 to adopt DeepSeek V4, Kimi K2.6, GLM-5, and Qwen3.5, citing 80-90% of frontier capability at 70x lower cost, per CNBC.
A year ago, Chinese open-weight models were about 4.5% of US enterprise API traffic. The current figure is 30 to 46%. A bank that answers to compliance officers reportedly told regulated clients to route real work to models built in China. That's a cost gap wide enough to override the reflex to buy American frontier.
I've been skeptical of the 'Chinese models caught up' story. The benchmarks felt cherry-picked and the deployment stories were thin. A formal recommendation from Goldman to a compliance-bound institution is different evidence than a leaderboard screenshot. Goldman isn't in the hype business.
Seventy times cheaper at 80-90% capability is the number that reorganizes a budget. Most production traffic isn't frontier-hard. You don't need GPT-5.6 to classify a support ticket or pull fields off an invoice. Paying frontier prices for that work was habit and vendor gravity, not necessity.
Stop treating 'which model' as a fixed decision. Build an eval suite against your real workload, drop DeepSeek V4 and Qwen3.5 into it, and measure. If they clear your quality bar on 70% of traffic, route that share. Keep frontier for the hard 30%. That's recurring savings, not a one-time win.
None of this settles the harder question. Running an open-weight model on your own hardware is one thing. Calling a hosted Chinese API with customer data is another, and export-control and data-residency rules haven't caught up to either. Read the license and know where the weights actually run before you wire one in.
Each link below shares sources, entities, or timing with this story.
The open-weight race just changed constraint. Moonshot AI suspended all new consumer subscriptions on July 20, roughly 48 hours after Kimi K3 launched, because request volume pushed its compute cluster to capacity. Remaining GPUs are reserved for existing paid subscribers. Tec...
DeepClaude hit 470 points on Hacker News. It swaps Claude Code's API backend to DeepSeek V4 Pro while preserving the full agent loop: file editing, bash execution, git tooling, the whole workflow. DeepSeek V4 Pro scores 96.4% on LiveCodeBench at a fraction of Anthropic's prici...
This is a supply-chain fact, and most people are still treating it as a geopolitics argument. Sequoia published "America's Open-Model Paradox" on July 24 with the number that reframes the whole conversation: Qwen's share of open-model fine-tunes went from 1% in January 2024 to...
Martin Alderson's essay "The upcoming AI margin collapse, part 1: GLM 5.2" hit 675 points and 462 comments on Hacker News, and it's the rare HN chart-topper that's actually about spreadsheet math instead of vibes. The argument is simple. Z.ai's GLM 5.2 delivers frontier-adjace...
The assumption that proprietary models own the coding benchmark crown just broke. Moonshot AI's Kimi K2.6 leads on 5 of 8 major agentic coding benchmarks while being the only open-weight model in the top tier. SWE-Bench Pro: 58.6% vs GPT-5.4's 57.7% and Claude Opus 4.6's 53.4%...
OpenAI cut GPT-5.6 Luna roughly 80%, from $1 to $0.20 per million input and $6 to $1.20 output. Anthropic priced Opus 5 at $5/$25 per million, half of Fable 5. The trigger is DeepSeek, Zhipu's GLM-5.2 and Moonshot's Kimi K3 landing 60-90% below US flagship pricing, with DoorDa...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.