Fetching from the wire…
Public story · 2026-08-25 · high
The pulled card listed 125B main-model parameters, 51B of n-gram embeddings, and 6B active per token before Qwen removed the details.
Why now: Qwen pulled the architecture details from the listing within minutes, so the screenshots commenters saved before the edit are the only public record of the original numbers.
Qwen listed Qwen3.8-Flash-Next on ModelScope with a model card describing it as an early preview of the company's next-generation Qwen4 architecture. Qwen edited out that description within minutes. Commenters had already saved screenshots of the original wording.
The card put real numbers on a Qwen4-class design for the first time. It described a redesigned multimodal mixture-of-experts model: 125B main-model parameters, an additional 51B in n-gram embeddings, and 6B active per token. Qwen said it was releasing the preview early so developers could start preparing for the architecture change.
The listing included an FP8 version. There was no FP4 build and no QAT-only release, so anyone planning to run this at lower precision doesn't have an official option yet.
The precedent here is Qwen3-Next, the last time Qwen introduced a new model architecture instead of an update to an existing one. That model took about two months to reach llama.cpp support after its release. The edited Qwen3.8-Flash-Next card didn't include a release date, just the preview.
Each link below shares sources, entities, or timing with this story.
Staged on ModelScope for 23:00 Beijing time August 26, it's roughly 125B parameters plus a separate N-gram embedding table of about 51B, activating 6B per token, with GDN gated-delta hybrid layers and Qwen Sparse Attention. Alibaba frames it as a technology preview of the Qwen...
The Hugging Face page is marked "Upcoming release" with no model card, license, architecture details, context length or benchmarks, after Alibaba promised both Qwen3.8-Max and the 27B weights for the week of August 10. A ModelScope countdown pointed at August 15. Unsloth signa...
Alibaba announced it August 3: sparse MoE with ~95B active per token, 1M context, 128k max output, $2/M input and $6/M output with $0.25/M cached. 67.4 on Terminal-Bench 2.1 (up from 61.0 for 3.7 Max), #4 on Frontend Code Arena at 1,668 Elo, #2 on Vals Index among open-weight...
Alibaba released Qwen3.6-27B on April 22. Dense architecture. Open weights. 77.2% on SWE-bench Verified, within 3.7 points of Claude Opus 4.6. On SkillsBench, it scores 48.2% versus its own 397B MoE predecessor's 30.0%. That's a 77% improvement with 14.8x fewer parameters. Let...
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's...
Hugging Face published its Summer 2026 State of Open Models report on August 14, and one statistic in it went almost entirely unremarked in the coverage. By July 2026, agents rather than humans became the Hub's primary users. Claude Code alone accounted for 44.4% of all agent...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.