Fetching from the wire…
Models2026-08-06 · source-backed
128K context, 2.6B parameters, posting 51.87 AIME25, 59.17 IFBench, 56.88 BFCLv4 and 62.85 Claw-Eval, beating Gemma-4-E4B (8B) across all four and trailing Qwen3.5-9B narrowly. 220 tok/s on an M5 Max CPU, ~30 tok/s on phones, day-one support in llama.cpp, MLX, vLLM, SGLang and ONNX. Local tool-calling agents just became genuinely viable on consumer hardware, which changes the calculus for anything privacy-sensitive.
Each link below shares sources, entities, or timing with this story.
Google released open-source Multi-Token Prediction (MTP) drafters for the Gemma 4 model family. The concept: pair a heavy target model (Gemma 4 31B) with a lightweight drafter that predicts several future tokens in parallel. The target model verifies the predictions in a singl...
For about a year, "run your agent locally" meant accepting a model that couldn't reliably call a tool twice in a row. That excuse is gone. Meta Superintelligence Labs published Muse Glimmer today: a 29.6B dense causal transformer, 52 layers, 6,656 hidden dim, with a ~1.8B ViT-...
The models are good. The license is the real story. Google released Gemma 4 on April 2 with four variants: E2B, E4B, 26B MoE, and 31B Dense. All built on the Gemini 3 architecture. The 31B Dense variant claimed #3 on Arena AI's text leaderboard, beating models 20x its size. Th...
Gemma 4 12B dropped June 3, and the spec sheet is the kind of thing I read twice to make sure I wasn't misreading it. 11.95 billion params, Apache 2.0, reads text, image, audio, and video. No separate vision encoder. No separate audio encoder. The model handles all of it nativ...
LFM2.5-DSpark is roughly 296M to 328M parameters, paired with LFM2.5 1.2B, 2.6B and 8B-A1B targets. Reported: up to 3.18x on GPU, 2.87x on-device, the 2.6B hitting 2.67x on an H100 and 2.27x on an M4 Max MacBook. Hugging Face The design combines a DFlash-style parallel backbon...
Google DeepMind released Gemma 4 on April 2 with four model sizes (E2B, E4B, 26B MoE, 31B Dense) under Apache 2.0. Multimodal (text, vision, audio). 256K context. Native thinking and tool-calling optimized for agentic workflows. Day-zero ecosystem support across vLLM, llama.cp...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.