Fetching from the wire…
Top 5 · 2026-04-11 · source-backed
Zhipu's GLM-5.1 ranked third on Code Arena, jumping 90+ points over its predecessor GLM-5 and landing ahead of GPT-5.4 and Gemini 3.1 Pro. Two separate r/LocalLLaMA threads (491 upvotes and 233 upvotes) confirm this isn't just a benchmark curiosity. Practitioners are paying attention.
The hard numbers: GLM-5.1 topped SWE-Bench Pro with 58.4 (GPT-5.4 got 57.7, Opus 4.6 got 57.3). That's the first time an open-source model has led that benchmark globally. It achieves 94.6% of Opus 4.6's coding performance at roughly one-third the cost.
I want to be careful here. Benchmarks aren't production. SWE-Bench Pro measures one slice of coding ability, and Code Arena rankings shift. But the pattern is hard to ignore. A year ago, open models were interesting for local inference and privacy-sensitive workloads. They weren't serious contenders for production coding agents. That's changing.
The second r/LocalLLaMA thread focused specifically on agentic benchmarks, where GLM-5.1 outperforms everything except Opus 4.6. For builders choosing base models for autonomous agent workloads, especially high-volume tasks where per-token cost matters, this shifts the cost-performance frontier meaningfully.
There's a bigger story here too. Apple's head of cloud told reporters that open-source models will address 90% of use cases. The r/LocalLLaMA community is voting on Qwen 3.6 feature priorities (592 upvotes, 260 comments). And meanwhile, DeepSeek has gone quiet (201 upvotes asking "what happened?"). The open-weight competitive map is reshuffling: Gemma 4, Qwen 3.5, and now GLM-5.1 are the ones to watch.
For builders: if you're running agent workloads where you're paying per token at scale, benchmark GLM-5.1 against your current model on your actual tasks. Not on SWE-Bench. On your codebase, your ticket types, your review standards. If it gets within 90% of your current quality at a third the price, the math does itself.
Each link below shares sources, entities, or timing with this story.
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's...
GLM-5.1 from Zhipu AI scored 58.4% on SWE-bench Pro. GPT-5.4 scored 57.7%. Claude Opus 4.6 scored 57.3%. That's the first time an open-weight model has ever topped a major coding benchmark against the best proprietary models. The specs matter. GLM-5.1 is a 754B-parameter mixtu...
An open-weight model just beat every closed frontier model on the benchmark builders actually care about. Z.AI (formerly Zhipu AI) dropped GLM-5.1, a 754-billion parameter mixture-of-experts model with 40 billion active parameters. The SWE-Bench Pro score: 58.4%. That's above...
I've spent the last year assuming that if I wanted real agentic coding quality, I paid for a closed model. That assumption took a hit on June 1. MiniMax shipped M3 with a new sparse-attention architecture (they call it MSA) that handles up to 1M tokens at roughly 9x prefill an...
GLM-5.1 scored 58.4% on SWE-Bench Pro. Opus 4.6 scored 57.3%. GPT-5.4 scored 57.7%. Read those numbers again. An open-weight, MIT-licensed model now leads the most rigorous coding benchmark we have. This isn't a narrow win on a cherry-picked eval. SWE-Bench Pro tests real-worl...
For about a year the ambient message has been "coding is mostly solved." SWE-Bench numbers crept past 50%, vendors put them on slides, and a lot of people quietly concluded the hard part was over. Cognition just dropped a bucket of cold water on that. On June 8 they launched F...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.