Fetching from the wire…
Models2026-08-27 · source-backed
The repo appeared August 24, opening the weights of a multimodal MoE that Qwen frames explicitly as an architecture preview, the same role Qwen3-Next played for Qwen3.5 (GitHub). The hybrid Gated DeltaNet plus Gated Attention design it previews already carried through the Qwen3.5 to Qwen3.8 series. Downstream corroboration arrived fast: ollama v0.33.1 lists MLX support and transformers v5.16.0 added the related Qwen4-Exp. A 235-upvote r/LocalLLaMA gallery on August 27 shows it beating DeepSeek V4 Pro on comparison benchmarks, with the practitioner caveat that llama.cpp support is still an unmerged PR, MTP doesn't work, and KV cache scaling is odd (r/LocalLLaMA). Running it locally this week means compiling a PR.
Each link below shares sources, entities, or timing with this story.
Staged on ModelScope for 23:00 Beijing time August 26, it's roughly 125B parameters plus a separate N-gram embedding table of about 51B, activating 6B per token, with GDN gated-delta hybrid layers and Qwen Sparse Attention. Alibaba frames it as a technology preview of the Qwen...
For about a year, "run your agent locally" meant accepting a model that couldn't reliably call a tool twice in a row. That excuse is gone. Meta Superintelligence Labs published Muse Glimmer today: a 29.6B dense causal transformer, 52 layers, 6,656 hidden dim, with a ~1.8B ViT-...
Somebody diffed the configs. Zero architectural changes. Same 64 layers, same 5,120 hidden dimension, same hybrid Gated DeltaNet → FFN / Gated Attention → FFN block structure as Qwen3.6-27B. The r/LocalLLaMA post showing this hit 945 upvotes and 157 comments, and Hugging Face...
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's...
Alibaba released Qwen3.6-27B on April 22. Dense architecture. Open weights. 77.2% on SWE-bench Verified, within 3.7 points of Claude Opus 4.6. On SkillsBench, it scores 48.2% versus its own 397B MoE predecessor's 30.0%. That's a 77% improvement with 14.8x fewer parameters. Let...
Alibaba released Qwen3.5-0.8B, 2B, 4B, and 9B — all natively multimodal (text+image+video from same weights, no adapter), 262K context, Apache 2.0. The 9B beats last-gen Qwen3-30B across the board and outperforms GPT-5-Nano by 13 points on MMMU-Pro. Architecture uses Gated Del...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.