Fetching from the wire…
Top 5 · 2026-04-09 · source-backed
GLM-5.1 scored 58.4% on SWE-Bench Pro. Opus 4.6 scored 57.3%. GPT-5.4 scored 57.7%. Read those numbers again. An open-weight, MIT-licensed model now leads the most rigorous coding benchmark we have.
This isn't a narrow win on a cherry-picked eval. SWE-Bench Pro tests real-world software engineering: resolving actual GitHub issues in real codebases. The model also demonstrated 8-hour autonomous execution capability, meaning it can work on a problem for an entire workday without human intervention. The 754B MoE architecture activates a fraction of parameters per inference, keeping it practical.
The economics are where this gets interesting for builders. GLM-5.1 via API costs $1.40 per million input tokens. Or you can run it yourself via vLLM or SGLang for the cost of your hardware. I run daily agent pipelines that make thousands of API calls. The cost difference between "free on my own GPU" and "$1.40/M tokens times a few thousand calls" is substantial over a month.
I've been skeptical of open-weight models catching frontier providers for coding tasks. The gap has been real. Llama models are great for many things, but complex multi-file code reasoning has been a consistent weakness. GLM-5.1 changes that calculation. Not by a little. By pulling ahead.
The timing creates an interesting dynamic. This lands the same week Anthropic launches managed agent infrastructure (Story #2) and the same week the community is loudly complaining about Opus 4.6 reasoning degradation (see Models below). If your agent pipeline depends on a single proprietary API, you now have a viable self-hosted alternative that benchmarks better on the thing that matters most: actually resolving engineering problems.
I'm not saying everyone should switch tomorrow. Benchmarks aren't everything, and I haven't personally run GLM-5.1 through my own workflows yet. But the argument for vendor lock-in just got a lot weaker. Every team running daily coding agents should evaluate GLM-5.1 as either a primary model or a fallback. The MIT license means you can modify, fine-tune, and deploy it however you want. No usage restrictions.
Something's shifted. The moat around proprietary coding models was already eroding. This might be the moment it disappeared.
Each link below shares sources, entities, or timing with this story.
GLM-5.1 from Zhipu AI scored 58.4% on SWE-bench Pro. GPT-5.4 scored 57.7%. Claude Opus 4.6 scored 57.3%. That's the first time an open-weight model has ever topped a major coding benchmark against the best proprietary models. The specs matter. GLM-5.1 is a 754B-parameter mixtu...
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's...
The assumption that proprietary models own the coding benchmark crown just broke. Moonshot AI's Kimi K2.6 leads on 5 of 8 major agentic coding benchmarks while being the only open-weight model in the top tier. SWE-Bench Pro: 58.6% vs GPT-5.4's 57.7% and Claude Opus 4.6's 53.4%...
Zhipu's GLM-5.1 ranked third on Code Arena, jumping 90+ points over its predecessor GLM-5 and landing ahead of GPT-5.4 and Gemini 3.1 Pro. Two separate r/LocalLLaMA threads (491 upvotes and 233 upvotes) confirm this isn't just a benchmark curiosity. Practitioners are paying at...
This is the other half of the Fable 5 story, so read them together. While the best coding model in the world is uncallable, an open-weight one quietly posted frontier-adjacent numbers. Per Tom's Hardware, independent benchmarks for the MIT-licensed GLM-5.2 (744B params, 40B ac...
Xiaomi released MiMo-V2.5-Pro, a 1.02 trillion parameter mixture-of-experts model (42B active) with 1M token context, fully MIT licensed. In benchmarks, it achieves 63.8% success on agentic tasks using 40-60% fewer tokens than Claude Opus 4.6 or GPT-5.4 for comparable results....
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.