Fetching from the wire…
Models2026-08-18 · source-backed
An r/LocalLLaMA thread asking where the promised MoE went turned up a hard artifact: modelscope/ms-swift commit a45f1d4, titled "fix wrong model-ids," removes the Qwen/Qwen3.8-35B-A3B and -FP8 entries from swift/model/models/qwen.py and substitutes the dense 27B. Why it matters for limited VRAM: Artificial Analysis has Qwen3.5-27B at 35, Qwen3.6-27B at 38, Qwen3.8-27B at 52, while the 3.6-generation 35B-A3B scored 32 at roughly 5x inference speed on 3B active params. A 3.8-generation sparse variant near that dense jump would be the best local-agent option of the year. The registry edit is the only concrete evidence and it points the wrong way.
Each link below shares sources, entities, or timing with this story.
Google DeepMind released Gemma 4 on April 2 with four model sizes (E2B, E4B, 26B MoE, 31B Dense) under Apache 2.0. Multimodal (text, vision, audio). 256K context. Native thinking and tool-calling optimized for agentic workflows. Day-zero ecosystem support across vLLM, llama.cp...
Alibaba released Qwen3.6-27B on April 22. Dense architecture. Open weights. 77.2% on SWE-bench Verified, within 3.7 points of Claude Opus 4.6. On SkillsBench, it scores 48.2% versus its own 397B MoE predecessor's 30.0%. That's a 77% improvement with 14.8x fewer parameters. Let...
The models are good. The license is the real story. Google released Gemma 4 on April 2 with four variants: E2B, E4B, 26B MoE, and 31B Dense. All built on the Gemini 3 architecture. The 31B Dense variant claimed #3 on Arena AI's text leaderboard, beating models 20x its size. Th...
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's...
Somebody diffed the configs. Zero architectural changes. Same 64 layers, same 5,120 hidden dimension, same hybrid Gated DeltaNet → FFN / Gated Attention → FFN block structure as Qwen3.6-27B. The r/LocalLLaMA post showing this hit 945 upvotes and 157 comments, and Hugging Face...
Alibaba published a fine-grained MoE with 2.4T total / 95B active, 512 experts, and a 92-layer hybrid full/linear attention backbone. vLLM shipped day-0 support verified on NVIDIA and AMD with ready 4-bit checkpoints (NVFP4 at 1.32 TiB for an 8xB300 node, MXFP4 at 1.45 TiB for...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.