Fetching from the wire…
Public story · 2026-08-17 · high
The August 16 release also lifts image LoRA limits to 10 and adds speculative decoding, four days after Qwen 3.8 shipped.
Why now: The release landed August 16, four days after Qwen 3.8 shipped, adding day-one support rather than catching up later.
KoboldCpp shipped version 1.119 on August 16, adding Minimax H3 video and image-to-video generation to its local binary, per r/LocalLLaMA. That folds video into the same tool people already run for local text and image models. Anyone doing everything on one GPU skips standing up a second runtime just for video.
The release also brings full Qwen 3.8 support, arriving four days after the model shipped, with jinja templates and tool calling built in. Muse Glimmer support arrived alongside it, and KoboldCpp added DSpark and Dflash speculative decoding to speed up generation.
Practical limits moved too. Max image and audio attachments per request rose to 64. Runtime image LoRAs went from 4 to 10, and Mistral's reasoning-budget controls now work inside the binary.
The announcement drew 63 upvotes, a number that undersells what shipped. A single local binary now handles text generation, tool calling, image LoRAs and video, all four days after a major model release. That's the kind of feature breadth that used to need three separate projects stitched together. The muted reception says local-model watchers are still grading releases on benchmark scores, not on how much a single binary can do.
Each link below shares sources, entities, or timing with this story.
For about a year, "run your agent locally" meant accepting a model that couldn't reliably call a tool twice in a row. That excuse is gone. Meta Superintelligence Labs published Muse Glimmer today: a 29.6B dense causal transformer, 52 layers, 6,656 hidden dim, with a ~1.8B ViT-...
Alibaba released Qwen3.6-27B on April 22. Dense architecture. Open weights. 77.2% on SWE-bench Verified, within 3.7 points of Claude Opus 4.6. On SkillsBench, it scores 48.2% versus its own 397B MoE predecessor's 30.0%. That's a 77% improvement with 14.8x fewer parameters. Let...
Somebody diffed the configs. Zero architectural changes. Same 64 layers, same 5,120 hidden dimension, same hybrid Gated DeltaNet → FFN / Gated Attention → FFN block structure as Qwen3.6-27B. The r/LocalLLaMA post showing this hit 945 upvotes and 157 comments, and Hugging Face...
The models are good. The license is the real story. Google released Gemma 4 on April 2 with four variants: E2B, E4B, 26B MoE, and 31B Dense. All built on the Gemini 3 architecture. The 31B Dense variant claimed #3 on Arena AI's text leaderboard, beating models 20x its size. Th...
Google DeepMind released Gemma 4 on April 2 with four model sizes (E2B, E4B, 26B MoE, 31B Dense) under Apache 2.0. Multimodal (text, vision, audio). 256K context. Native thinking and tool-calling optimized for agentic workflows. Day-zero ecosystem support across vLLM, llama.cp...
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.