Fetching from the wire…
Top 5 · 2026-06-21 · source-backed
This is the other half of the Fable 5 story, so read them together. While the best coding model in the world is uncallable, an open-weight one quietly posted frontier-adjacent numbers. Per Tom's Hardware, independent benchmarks for the MIT-licensed GLM-5.2 (744B params, 40B active MoE, 1M-token context) landed: 81.0 on Terminal-Bench 2.1, within 4 points of Claude Opus 4.8's 85.0, and 62.1 on SWE-bench Pro, ahead of GPT-5.5's 58.6.
Let me be specific about what "within 4 points" means in practice, because benchmark proximity usually lies. Four points on Terminal-Bench is the difference between a model that mostly finishes multi-step terminal tasks and one that mostly finishes them with one more retry. It's not nothing. But it's the gap between "uncallable and #1" and "callable and very good," and right now callable wins every time.
The licensing is the headline for builders. MIT, self-hostable, with day-one support for Claude Code, Cline, and Roo Code. You can run the thing on your own GPUs and nobody can revoke it. Sebastian Raschka's June 18 Ahead of AI note dissects the IndexShare mechanism that makes its 1M-token sparse attention cheap enough to self-host. That's the architectural "why" behind the pricing, and it's worth reading before you commit GPU budget, because long-context inference cost is where self-hosting projects quietly die.
On June 19-20, Z.ai co-founder Jie Tang publicly rebutted Elon Musk's claim that China is a year out from a Fable-5-class model. Tang's words: "it won't take that long." Given that Fable 5 is currently banned and GLM-5.2 is four points behind Opus 4.8 and shipping today, I'd say Tang is winning that argument by default.
What builders should do: stand up a GLM-5.2 endpoint as a fallback now, even if it's not your primary. The displaced Fable 5 demand is flowing to GLM-5.2 and GPT-5.5, and capacity is the constraint. Get your config tested while it's calm, not during the next outage. I wired a fallback route last week and the swap took an afternoon. Worth it.
Each link below shares sources, entities, or timing with this story.
Everyone spent yesterday arguing about benchmark numbers. Tencent quietly published data suggesting the numbers belong to your infrastructure, not the model. The WorkBuddy Bench leaderboard reports every model under two different agent harnesses — CodeBuddy Code and Claude Cod...
Z.ai's GLM-5 (744B/40B MoE, MIT license, 205K context) is free on NVIDIA NIM at 40 req/min with no credit card. Benchmarks: 77.8% SWE-bench Verified (highest open-source), 56.2 Terminal-Bench 2.0 (approaching Opus 4.5's 59.3). Trained entirely on 100,000 Huawei Ascend chips. Y...
Martin Alderson's essay "The upcoming AI margin collapse, part 1: GLM 5.2" hit 675 points and 462 comments on Hacker News, and it's the rare HN chart-topper that's actually about spreadsheet math instead of vibes. The argument is simple. Z.ai's GLM 5.2 delivers frontier-adjace...
Weeks after launch, Z.ai's open-weight GLM-5.2 now accounts for roughly 75% of all Z.ai model traffic on OpenRouter, with at least one provider serving it past 125 tokens per second (GIGAZINE, citing OpenRouter). The numbers behind the surge: an Artificial Analysis Intelligenc...
GLM-5.1 from Zhipu AI scored 58.4% on SWE-bench Pro. GPT-5.4 scored 57.7%. Claude Opus 4.6 scored 57.3%. That's the first time an open-weight model has ever topped a major coding benchmark against the best proprietary models. The specs matter. GLM-5.1 is a 754B-parameter mixtu...
The 397B MoE scores 86.1 on Terminal-Bench 2.1 (Terminus-2) against Claude Opus 4.8's 85.0, 86.0 on SWE-bench Verified, and 92.8 on GPQA Diamond. Hugging Face But it trails badly on the harder agentic rows: 13.5 versus 21.1 on Frontier-Bench v0.1, 59.5 versus 69.7 on NL2Repo....
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.