Fetching from the wire…
Source-backed findings, relationship evidence, citations, and briefing history from the public MindPattern archive.
Qwen3.8-Flash-Next reaches 25 tok/s on M2 Ultra
Source findingQwen3.8-Flash-Next reaches 45 tok/s on M4 Max
Source findingslotstream runs Qwen3.8-Flash-Next model on macOS by streaming experts off SSD.
Source findingUnsloth enables multi-token prediction by default for Qwen3.8-Flash-Next
Source findingQwen3.8-Flash-Next runs in SGLang with SSD-streamed n-gram table optimization.
Source findingQwen3.8-Flash-Next leads Hugging Face trending with 3x score of GLM-5.3-Flash
Source findingPer-layer 97.7 GiB n-gram embeddings table can be disk-resident
Source findingPR 27742 merged Qwen3.8-Flash-Next with memory-mapped n-gram table support
Source findingUnsloth beta added local support for Qwen3.8-Flash-Next
Source findingAlibaba is shipping Qwen3.8-Flash-Next as a technology preview of Qwen4 architecture.
Source findingUnsloth released day-0 quantizations for Qwen3.8-Flash-Next.
Source findingQwen3.8-Flash-Next was released on Hugging Face with estimated drop on 2026-08-26.
Source finding