Fetching from the wire…
Top 5 · 2026-04-02 · source-backed
The models are good. The license is the real story.
Google released Gemma 4 on April 2 with four variants: E2B, E4B, 26B MoE, and 31B Dense. All built on the Gemini 3 architecture. The 31B Dense variant claimed #3 on Arena AI's text leaderboard, beating models 20x its size. The 26B MoE sits at #6. Multimodal input (video and images), 128K to 256K context windows, 140+ languages. Available today on HuggingFace, Ollama, Kaggle, and Google AI Studio.
Those are strong numbers. But I've seen strong numbers before from Gemma. What I haven't seen is Apache 2.0.
Previous Gemma licenses were restrictive enough to block enterprise adoption. You couldn't use them for competitive model training. Commercial deployments had legal gray areas. That's gone. Apache 2.0 means you can fine-tune, distill, embed, and ship Gemma 4 in your product with the same freedom you'd have with Llama or Mistral. For the local LLM community, this is the moment Google stops being the walled-garden option and starts competing directly with Meta's open model strategy.
I've been watching r/LocalLLaMA for the initial benchmarks, and the early reports are strong on coding and reasoning tasks specifically. The E2B and E4B variants cover edge deployment, while the 26B MoE and 31B Dense cover server-side inference. Four points on the size-capability curve from one release, all Apache 2.0, all with multimodal input. That's a complete lineup, not a single model launch.
The timing matters too. Vitalik Buterin published a self-sovereign local LLM guide the same day, testing Qwen3.5:35B locally and declaring 2026 "the year to reclaim computing self-sovereignty." Two independent signals converging on local-first AI on the same date. Gemma 4 at Apache 2.0 is the kind of model that makes that vision practical.
If you're evaluating open models for any production workload, benchmark Gemma 4 against Qwen and Llama today. The Apache 2.0 licensing alone might make it your default choice.
Each link below shares sources, entities, or timing with this story.
Google DeepMind released Gemma 4 on April 2 with four model sizes (E2B, E4B, 26B MoE, 31B Dense) under Apache 2.0. Multimodal (text, vision, audio). 256K context. Native thinking and tool-calling optimized for agentic workflows. Day-zero ecosystem support across vLLM, llama.cp...
Google released Gemma 4 on April 2 with four model variants: E2B, E4B, 26B MoE, and 31B Dense. The license change is the first thing worth noting. Every previous Gemma had restrictions that made lawyers nervous. Gemma 4 is Apache 2.0. Full stop. Use it in any product, any way...
For about a year, "run your agent locally" meant accepting a model that couldn't reliably call a tool twice in a row. That excuse is gone. Meta Superintelligence Labs published Muse Glimmer today: a 29.6B dense causal transformer, 52 layers, 6,656 hidden dim, with a ~1.8B ViT-...
Xiaomi released MiMo-V2.5-Pro, a 1.02 trillion parameter mixture-of-experts model (42B active) with 1M token context, fully MIT licensed. In benchmarks, it achieves 63.8% success on agentic tasks using 40-60% fewer tokens than Claude Opus 4.6 or GPT-5.4 for comparable results....
Alibaba released Qwen3.6-27B on April 22. Dense architecture. Open weights. 77.2% on SWE-bench Verified, within 3.7 points of Claude Opus 4.6. On SkillsBench, it scores 48.2% versus its own 397B MoE predecessor's 30.0%. That's a 77% improvement with 14.8x fewer parameters. Let...
Google dropped Gemma 4 and it's not incremental. The 31B dense model ranks #3 on Arena AI with an ELO of 1,452, scores 85.2% on MMLU Pro, 89.2% on AIME 2026, and 80.0% on LiveCodeBench v6. It outperforms models 20x its size. Under Apache 2.0. At $0.20 per run. Only Opus 4.6 an...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.