Fetching from the wire…
Public story · 2026-08-05 · high
Maple-Preview hits 218 tokens per second on a Mac mini M4, 5 to 16 times faster than Gemma 4 and Qwen3.5, DeepGrove says.
Why now: On August 5, Maple-Preview's traction comes from a Show HN thread past 140 points, not from any name recognition DeepGrove has yet.
DeepGrove published Maple-Preview, a 20-billion-parameter mixture-of-experts model that runs with just 1 billion parameters active per token, in a 5.31GB checkpoint, per its Hugging Face listing.
On a Mac mini M4, DeepGrove clocks it at 218 tokens per second. That's fast enough to run a 20B model locally without a GPU server or an API bill. That's 5 to 16 times faster than Gemma 4, Qwen3.5 and gpt-oss at comparable quality, DeepGrove reports. It backs that with scores on four benchmarks: LCBv6, AIME 2026, HMMT 2026 and GPQA-D.
The model has 24 layers and 256 experts, with 8 active on any given pass. Context runs to 131,072 tokens, built from a 3:1 mix of sliding-window and global attention.
The weights are ternary, which shrinks the model itself instead of paging a large one off disk. That's a different bet than the SSD-streaming approach other small-model projects use. There, a large model stays on disk and pages sections into memory as needed.
The Show HN thread backing the release passed 140 points, with one commenter reporting 120 tokens per second on an iPhone. It ships under MIT, so anyone can test the speed and quality claims independently.
Each link below shares sources, entities, or timing with this story.
Google's HF org lists diffusiongemma-26B-A4B-it (~4B active), an image-text-to-text Gemma member that's diffusion-style rather than purely autoregressive (Hugging Face). No detailed announcement yet, which is why I'm flagging it low. But a diffusion approach inside the Gemma o...
Hugging Face published its Summer 2026 State of Open Models report on August 14, and one statistic in it went almost entirely unremarked in the coverage. By July 2026, agents rather than humans became the Hub's primary users. Claude Code alone accounted for 44.4% of all agent...
For about a year, "run your agent locally" meant accepting a model that couldn't reliably call a tool twice in a row. That excuse is gone. Meta Superintelligence Labs published Muse Glimmer today: a 29.6B dense causal transformer, 52 layers, 6,656 hidden dim, with a ~1.8B ViT-...
Google released open-source Multi-Token Prediction (MTP) drafters for the Gemma 4 model family. The concept: pair a heavy target model (Gemma 4 31B) with a lightweight drafter that predicts several future tokens in parallel. The target model verifies the predictions in a singl...
Xiaomi released MiMo-V2.5-Pro, a 1.02 trillion parameter mixture-of-experts model (42B active) with 1M token context, fully MIT licensed. In benchmarks, it achieves 63.8% success on agentic tasks using 40-60% fewer tokens than Claude Opus 4.6 or GPT-5.4 for comparable results....
The models are good. The license is the real story. Google released Gemma 4 on April 2 with four variants: E2B, E4B, 26B MoE, and 31B Dense. All built on the Gemini 3 architecture. The 31B Dense variant claimed #3 on Arena AI's text leaderboard, beating models 20x its size. Th...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.