Fetching from the wire…
Public story · 2026-08-10 · high
The 29.6B model quantizes to under 20GB, runs on 24 to 32GB of consumer hardware, and scores 51.2% on SWE-Bench Pro.
Why now: Muse Glimmer launched August 10, the same day Wang teased Spark 1.2 and a rival benchmark reignited the local-versus-hosted debate.
Meta Superintelligence Labs released Muse Glimmer on August 10, a 29.6 billion-parameter model for local AI agents licensed under Apache 2.0, per Meta's technical blog.
For about a year, local models couldn't reliably call a tool twice in a row. Muse Glimmer scores 51.2% on SWE-Bench Pro, the first credible open substitute for a hosted agent loop.
Quantized to roughly 4-bit, the model fits under 20GB and runs on 24 to 32GB of consumer hardware. AMD published a same-day guide for Ryzen AI Max and Radeon.
It was distilled from April's Muse Spark and built for always-on agents, naming OpenClaw among the supported orchestrators. Meta benchmarks it against Gemma4-31B and Qwen3.6-27B.
Open-source tooling outran Meta's own release. An Unsloth GGUF appeared on Hugging Face within hours, while Meta's post lists Ollama, LM Studio, llama.cpp, MLX, ExecuTorch, vLLM and SGLang as "coming soon." One Reddit commenter warned that template bugs cause most "can't tool-call" complaints.
DFlash, Meta's speculative-decoding drafter, proposes whole blocks of tokens for the base model to verify. That's 3.1x faster on an RTX 5090, but only 1.8x on M5 Max and 1.5x on M4 Max, per Meta's blog.
That "3x faster" framing assumes Nvidia hardware; on last-gen Apple silicon you get roughly half the CUDA benefit. I'm on a Mac, so that gap decides whether this replaces my hosted loop or just backs it up.
Alexandr Wang said on X that Meta will open-weight Muse Spark 1.2 "soon." That's the 1M-context frontier model that launched August 5, per the Wall Street Journal's "coming weeks" timeline.
An independent harness measured DeepSeek-V4-Flash-0731 at 82.7% on Terminal-Bench 2.1 that same day, up from 61.8% for the preview, at $0.14 per million input tokens. HN commenters used the price gap to argue against buying inference hardware at all.
Pull the Unsloth GGUF, check the chat template against your tool schema, then see if it holds as a free fallback for rate limits.
Each link below shares sources, entities, or timing with this story.
DeepSeek released V4 Preview on April 24 with two open-weight variants: V4-Pro (1.6T total parameters, 49B activated via MoE) and V4-Flash (284B parameters, 13B activated). Both support 1M-token context windows. Both are Apache 2.0 licensed. Both are live right now on Hugging...
Alibaba released Qwen3.6-27B on April 22. Dense architecture. Open weights. 77.2% on SWE-bench Verified, within 3.7 points of Claude Opus 4.6. On SkillsBench, it scores 48.2% versus its own 397B MoE predecessor's 30.0%. That's a 77% improvement with 14.8x fewer parameters. Let...
DeepSeek-V4-Flash-0731 landed July 31 under MIT with a DSpark speculative-decoding module attached. Terminal Bench 2.1: 82.7. Toolathlon-Verified: 70.3. DSBench-FullStack: 68.7. DeepSWE: 54.4. NL2Repo: 54.2. The model card claims it beats DeepSeek-V4-Pro (Preview) "despite its...
Three moves, two days, no coordination between them. August 10–11: GitHub shipped Ollama as a BYOK provider inside Copilot for JetBrains (GitHub Changelog). Unsloth released Unsloth Desktop with a command literally named unsloth start claude, which points Claude Code and Codex...
The models are good. The license is the real story. Google released Gemma 4 on April 2 with four variants: E2B, E4B, 26B MoE, and 31B Dense. All built on the Gemini 3 architecture. The 31B Dense variant claimed #3 on Arena AI's text leaderboard, beating models 20x its size. Th...
DeepSeek released V4 on April 24 and the numbers demand attention. V4-Pro is 1.6 trillion parameters total with 49 billion active, MIT-licensed, native 1M-token context. It scores 80.6% on SWE-bench Verified, putting it within 0.2 points of Claude Opus 4.6. On Terminal-Bench 2...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.