Fetching from the wire…
Top 5 · 2026-04-21 · source-backed
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's 53.4. On SWE-Bench Verified it hits 80.2. The 659-point HN thread has been running all day.
What makes this different from the usual benchmark chest-thumping is the agent swarm architecture. K2.6 scales to 300 concurrent sub-agents across 4,000 coordinated steps. Moonshot's documented tests include generating 100 tailored resumes, 40-page research papers, and 30 landing pages in single autonomous runs. This isn't autocomplete. It's a self-hosted agentic workforce.
The timing matters. On r/LocalLLaMA, an Opus 4.7 Max subscriber posted that they're migrating their entire team to K2.6. The post got 122 upvotes, and the comments are telling. The poster explicitly says they're not anti-Anthropic. They just found that open weights at this capability level change the math.
And it's not just K2.6. Alibaba released Qwen3.6-Max-Preview the same day, ranking first on SWE-bench Pro, Terminal-Bench 2.0, SkillsBench, and three other coding benchmarks. Two Chinese labs, same 24-hour window, both topping agentic coding leaderboards. Something's happening.
On the Artificial Analysis Intelligence Index, K2.6 at 54 sits just three points behind the closed-model trio at 57. That's the smallest frontier gap ever measured for open weights. Combined with Qwen's results and Gemma 4's efficient MoE architecture, the argument that frontier capability requires closed models is getting harder to make.
For builders: if you're running agentic pipelines and paying per-token for closed models, this week is when you should start benchmarking K2.6 against your actual workloads. Not on academic tasks. On your codebase, your ticket backlog, your deployment scripts. The model weights are available. Moonshot also updated their K2 Vendor Verifier to compare tool-call accuracy across 12 inference providers, because what you get from Provider A versus Provider B can differ meaningfully even with the same weights.
I don't know if K2.6 holds up across all the rough edges of real production work. Benchmarks are benchmarks. But a 1T open-weights model beating GPT-5.4 on the hardest coding benchmark while running 300 parallel agents is the kind of thing that shifts how you think about architecture.
Each link below shares sources, entities, or timing with this story.
Alibaba released Qwen3.6-27B on April 22. Dense architecture. Open weights. 77.2% on SWE-bench Verified, within 3.7 points of Claude Opus 4.6. On SkillsBench, it scores 48.2% versus its own 397B MoE predecessor's 30.0%. That's a 77% improvement with 14.8x fewer parameters. Let...
Zhipu's GLM-5.1 ranked third on Code Arena, jumping 90+ points over its predecessor GLM-5 and landing ahead of GPT-5.4 and Gemini 3.1 Pro. Two separate r/LocalLLaMA threads (491 upvotes and 233 upvotes) confirm this isn't just a benchmark curiosity. Practitioners are paying at...
The assumption that proprietary models own the coding benchmark crown just broke. Moonshot AI's Kimi K2.6 leads on 5 of 8 major agentic coding benchmarks while being the only open-weight model in the top tier. SWE-Bench Pro: 58.6% vs GPT-5.4's 57.7% and Claude Opus 4.6's 53.4%...
Xiaomi released MiMo-V2.5-Pro, a 1.02 trillion parameter mixture-of-experts model (42B active) with 1M token context, fully MIT licensed. In benchmarks, it achieves 63.8% success on agentic tasks using 40-60% fewer tokens than Claude Opus 4.6 or GPT-5.4 for comparable results....
GLM-5.1 from Zhipu AI scored 58.4% on SWE-bench Pro. GPT-5.4 scored 57.7%. Claude Opus 4.6 scored 57.3%. That's the first time an open-weight model has ever topped a major coding benchmark against the best proprietary models. The specs matter. GLM-5.1 is a 754B-parameter mixtu...
Weeks after launch, Z.ai's open-weight GLM-5.2 now accounts for roughly 75% of all Z.ai model traffic on OpenRouter, with at least one provider serving it past 125 tokens per second (GIGAZINE, citing OpenRouter). The numbers behind the surge: an Artificial Analysis Intelligenc...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.