Fetching from the wire…
Top 5 · 2026-05-01 · source-backed
Xiaomi released MiMo-V2.5-Pro, a 1.02 trillion parameter mixture-of-experts model (42B active) with 1M token context, fully MIT licensed. In benchmarks, it achieves 63.8% success on agentic tasks using 40-60% fewer tokens than Claude Opus 4.6 or GPT-5.4 for comparable results. In a demo, it built a complete SysY compiler in Rust in 4.3 hours with 672 tool calls.
This is the most capable fully open-source agentic model released to date. Full stop.
The timing matters. r/LocalLLaMA is calling April 2026 the "best month of all time" for local models. Six organizations shipped competitive open weights in a single month: Google (Gemma 4), Alibaba (Qwen 3.6), Meta (Llama 4), Mistral (Medium 3.5), Zhipu AI (GLM-5.1 744B MoE), and DeepSeek (V4). The gap between open and closed is collapsing faster than anyone predicted.
What catches my attention about MiMo isn't the parameter count. It's the token efficiency. Using 40-60% fewer tokens than frontier closed models for comparable agentic results means dramatically lower inference costs for self-hosted deployments. Combined with the MIT license, this opens the door for companies that can't or won't send proprietary code to Anthropic or OpenAI. And the 1M context window means you're not making compromises on what the model can hold in working memory.
For builders evaluating self-hosted agentic models, benchmark MiMo-V2.5-Pro against your current setup this week. If you're paying per-token for agentic workflows and the quality holds, the cost savings alone could justify the migration. If you're in a regulated industry where data can't leave your infrastructure, this might be the first open model that's actually good enough for production agent work.
One caveat: I haven't run it myself yet. The benchmarks look strong but benchmarks lie, especially for agentic tasks where real-world reliability matters more than peak performance. Test it on your actual workflows before committing.
Each link below shares sources, entities, or timing with this story.
GLM-5.1 from Zhipu AI scored 58.4% on SWE-bench Pro. GPT-5.4 scored 57.7%. Claude Opus 4.6 scored 57.3%. That's the first time an open-weight model has ever topped a major coding benchmark against the best proprietary models. The specs matter. GLM-5.1 is a 754B-parameter mixtu...
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's...
The models are good. The license is the real story. Google released Gemma 4 on April 2 with four variants: E2B, E4B, 26B MoE, and 31B Dense. All built on the Gemini 3 architecture. The 31B Dense variant claimed #3 on Arena AI's text leaderboard, beating models 20x its size. Th...
Moonshot AI released Kimi K3, a sparse mixture-of-experts activating 16 of 896 experts per token. That's about 1.8% of the pool live at any moment, with a 1M-token context window and native vision. Two new architectural pieces show up: Kimi Delta Attention and Attention Residu...
For about a year, "run your agent locally" meant accepting a model that couldn't reliably call a tool twice in a row. That excuse is gone. Meta Superintelligence Labs published Muse Glimmer today: a 29.6B dense causal transformer, 52 layers, 6,656 hidden dim, with a ~1.8B ViT-...
This is a supply-chain fact, and most people are still treating it as a geopolitics argument. Sequoia published "America's Open-Model Paradox" on July 24 with the number that reframes the whole conversation: Qwen's share of open-model fine-tunes went from 1% in January 2024 to...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.