Fetching from the wire…
Public story · 2026-08-26 · high
It activates 6B of roughly 125B parameters and claims a training cost one-ninth of Qwen3.7-Plus.
Why now: Weights are staged for release at 23:00 Beijing time on August 26, with r/LocalLLaMA and Unsloth already positioned to test it the moment it posts.
Alibaba staged Qwen3.8-Flash-Next on ModelScope for release at 23:00 Beijing time on August 26, calling it an early preview of the Qwen4 architecture.
The model activates just 6 billion of its roughly 125 billion parameters per token. Alibaba says training it cost about a ninth of Qwen3.7-Plus for similar capability. That ratio would reshape the budget for anyone deciding whether to train a model from scratch.
Qwen3.8-Flash-Next pairs those parameters with an N-gram embedding table of roughly 51 billion, GDN gated-delta hybrid layers, and Qwen Sparse Attention. Alibaba hasn't published any side-by-side scores against its own models or anyone else's. Treat the training-cost number as a claim, not a result.
r/LocalLLaMA already has a pinned megathread, and Unsloth announced day-0 quantized versions before the weights even existed. The top comment came from someone who'd just finished tuning Qwen3.8-27B, joking that the paint was still wet on that release. Commenters are asking for llama.cpp flags to push the sparse KV cache onto SSD.
Each link below shares sources, entities, or timing with this story.
The card described a redesigned multimodal MoE with 125B main-model parameters, an additional 51B of n-gram embeddings, and 6B active per token, stating it's built on the next-generation Qwen4 architecture and released early so the community can prepare (ModelScope). Commenter...
The repo appeared August 24, opening the weights of a multimodal MoE that Qwen frames explicitly as an architecture preview, the same role Qwen3-Next played for Qwen3.5 (GitHub). The hybrid Gated DeltaNet plus Gated Attention design it previews already carried through the Qwen...
The Hugging Face page is marked "Upcoming release" with no model card, license, architecture details, context length or benchmarks, after Alibaba promised both Qwen3.8-Max and the 27B weights for the week of August 10. A ModelScope countdown pointed at August 15. Unsloth signa...
Created August 24, it holds a trendingScore of 3,967 against second-place GLM-5.3-Flash at 1,376 (Hugging Face). The near-1:1 like-to-download ratio means almost everyone bookmarking it hasn't pulled weights, and the unsloth GGUF conversion at 4,354 downloads is absorbing comp...
Somebody diffed the configs. Zero architectural changes. Same 64 layers, same 5,120 hidden dimension, same hybrid Gated DeltaNet → FFN / Gated Attention → FFN block structure as Qwen3.6-27B. The r/LocalLLaMA post showing this hit 945 upvotes and 157 comments, and Hugging Face...
Alibaba released Qwen3.6-27B on April 22. Dense architecture. Open weights. 77.2% on SWE-bench Verified, within 3.7 points of Claude Opus 4.6. On SkillsBench, it scores 48.2% versus its own 397B MoE predecessor's 30.0%. That's a 77% improvement with 14.8x fewer parameters. Let...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.