Fetching from the wire…
Top 5 · 2026-04-06 · source-backed
Google dropped Gemma 4 and it's not incremental. The 31B dense model ranks #3 on Arena AI with an ELO of 1,452, scores 85.2% on MMLU Pro, 89.2% on AIME 2026, and 80.0% on LiveCodeBench v6. It outperforms models 20x its size. Under Apache 2.0. At $0.20 per run.
Only Opus 4.6 and GPT-5.2 beat it. Let that sink in.
The ecosystem response has been immediate and broad. Google's AI Edge Gallery app hit #8 on the App Store productivity charts. It runs Gemma 4 models entirely on-device, no internet required, under 1.5GB of memory. A developer benchmarked the 26B MoE variant on a MacBook Pro M4 Pro running LM Studio's new headless CLI and got 51 tokens per second. That's usable. Someone built PokeClaw, a working app that uses Gemma 4 to autonomously control an Android phone. No server, no cloud, no API calls. Just a phone running a model that rivals frontier systems.
The technical explanation for why it punches this far above its weight comes down to per-layer embeddings, an architecture where the 26B MoE variant only activates 3.8B parameters per pass. It's a 448-upvote technical explainer on r/LocalLLaMA and the clearest community breakdown of how Google pulled this off.
Here's what I'd actually do this week: take your three most expensive API-dependent features, benchmark them against Gemma 4 running locally, and calculate the savings. If you're spending real money on Sonnet calls for tasks that don't require Opus-level reasoning, Gemma 4 at $0.20/run or free on local hardware might just be the answer. The economics have shifted. Not theoretically. Right now.
Each link below shares sources, entities, or timing with this story.
The models are good. The license is the real story. Google released Gemma 4 on April 2 with four variants: E2B, E4B, 26B MoE, and 31B Dense. All built on the Gemini 3 architecture. The 31B Dense variant claimed #3 on Arena AI's text leaderboard, beating models 20x its size. Th...
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's...
Google DeepMind released Gemma 4 on April 2 with four model sizes (E2B, E4B, 26B MoE, 31B Dense) under Apache 2.0. Multimodal (text, vision, audio). 256K context. Native thinking and tool-calling optimized for agentic workflows. Day-zero ecosystem support across vLLM, llama.cp...
Xiaomi released MiMo-V2.5-Pro, a 1.02 trillion parameter mixture-of-experts model (42B active) with 1M token context, fully MIT licensed. In benchmarks, it achieves 63.8% success on agentic tasks using 40-60% fewer tokens than Claude Opus 4.6 or GPT-5.4 for comparable results....
For about a year, "run your agent locally" meant accepting a model that couldn't reliably call a tool twice in a row. That excuse is gone. Meta Superintelligence Labs published Muse Glimmer today: a 29.6B dense causal transformer, 52 layers, 6,656 hidden dim, with a ~1.8B ViT-...
GLM-5.1 from Zhipu AI scored 58.4% on SWE-bench Pro. GPT-5.4 scored 57.7%. Claude Opus 4.6 scored 57.3%. That's the first time an open-weight model has ever topped a major coding benchmark against the best proprietary models. The specs matter. GLM-5.1 is a 754B-parameter mixtu...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.