Fetching from the wire…
Top 5 · 2026-06-25 · source-backed
Weeks after launch, Z.ai's open-weight GLM-5.2 now accounts for roughly 75% of all Z.ai model traffic on OpenRouter, with at least one provider serving it past 125 tokens per second (GIGAZINE, citing OpenRouter). The numbers behind the surge: an Artificial Analysis Intelligence Index of 51, the first open model past 80% on Terminal-Bench 2.1, and 62.1 on SWE-bench Pro. Output costs run roughly 5 to 8 times cheaper than Claude Opus 4.8 and about one-sixth of GPT-5.5 Pro. The June 13 suspension of Claude Fable 5 and Mythos 5 poured gasoline on it. People had agentic coding workloads running, the model they were using went away, and GLM-5.2 was sitting right there at a quarter of the price.
Read this next to the Qwen story and the arc is hard to miss. Lab A allegedly harvests agentic-reasoning behavior. Open-weight Lab B ships a model that does agentic coding at 62 on SWE-bench Pro for pennies. I'm not claiming a direct line between those two facts. But the macro pattern is the same: the agentic-coding capability that was a frontier-lab exclusive eighteen months ago is now a commodity you route to by price and latency.
I've started doing this myself in personal projects. Not everything needs Opus. A lot of my agent work is mechanical: write the test, run it, read the failure, patch, repeat. For that loop, a model at 125 TPS and one-sixth the cost changes the math on how aggressively I can fan out. When each subagent is cheap, "spawn 30 of them" stops being a budget decision.
The catch, and I want to be honest here, is that I don't trust leaderboard SWE-bench numbers to predict how a model behaves on my codebase. Benchmark 62 and "good on a 40-file TypeScript repo with weird internal conventions" are different claims. So the move isn't to rip out Opus. It's to set up a real eval on your own traces (see the pass^k story below and the skills section) and let the data decide which workloads drop down to GLM-5.2. Route by evidence, not by leaderboard.
Each link below shares sources, entities, or timing with this story.
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's...
Everyone spent yesterday arguing about benchmark numbers. Tencent quietly published data suggesting the numbers belong to your infrastructure, not the model. The WorkBuddy Bench leaderboard reports every model under two different agent harnesses — CodeBuddy Code and Claude Cod...
This is the other half of the Fable 5 story, so read them together. While the best coding model in the world is uncallable, an open-weight one quietly posted frontier-adjacent numbers. Per Tom's Hardware, independent benchmarks for the MIT-licensed GLM-5.2 (744B params, 40B ac...
Terminal-Bench 2.1 results (entries dated June 17) put Codex CLI on GPT-5.5 first at 83.4%, Claude Code on Fable 5 second at 83.1%, and Claude Code on Opus 4.8 at 78.9%. The asterisk matters more than the ranking: Fable 5 and Mythos 5 have been export-suspended since June 12,...
I don't care that Grok 4.5 ranks #4. I care that it resolves a SWE-Bench Pro task with an average of 15,954 output tokens where Opus 4.8 spends 67,020. That's a 4.2x efficiency gap, and it lands straight in my monthly bill. SpaceXAI launched Grok 4.5 on July 8, a roughly 1.5T-...
Martin Alderson's essay "The upcoming AI margin collapse, part 1: GLM 5.2" hit 675 points and 462 comments on Hacker News, and it's the rare HN chart-topper that's actually about spreadsheet math instead of vibes. The argument is simple. Z.ai's GLM 5.2 delivers frontier-adjace...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.