Fetching from the wire…
Models2026-08-30 · source-backed
The AngelSlim/Hy4-preview-GGUF repo offers Q4_K_M at 435.20 GiB (4.86 bpw) and STQ1_0 at 213.66 GiB (2.38 bpw), benchmarked at 204.56 t/s prefill and 20.47 t/s decode on 8x H20. STQ1_0 comes from llama.cpp PR #22836 and uses ternary weights with 3:4 forced sparsity at 1.3125 bpw on routed-expert gate/up projections across 29 layers, with IQ2_XXS at 2.0625 bpw on the other 48. An imatrix is mandatory. (Hugging Face) The r/LocalLLaMA thread (821 upvotes) circulated a ~98% performance retention claim, and the top skeptical reply notes 98% on KL divergence alone can be misleading. I'd want task benchmarks before believing that number.
Each link below shares sources, entities, or timing with this story.
78 layers where layer one is dense FFN and the other 77 are MoE, each with 256 routed experts and 1 shared expert, top-8 routing per token, plus a native 10B MTP layer (0.7B activated) built in for speculative decoding. FP8 and base variants released together on August 28; the...
DeepSeek released V4 Preview on April 24 with two open-weight variants: V4-Pro (1.6T total parameters, 49B activated via MoE) and V4-Flash (284B parameters, 13B activated). Both support 1M-token context windows. Both are Apache 2.0 licensed. Both are live right now on Hugging...
On July 14, llama.cpp merged native support for Tencent's Hunyuan Hy3 architecture (PR #25395), a 295B-parameter, 21B-active MoE. Any recent master build can load it now. Community GGUF quants (Q2_K, IQ2_M, Q4_K_M) from AngelSlim and others already ship on Hugging Face, and so...
Tencent published UI-Mate-27B on Hugging Face, built on Qwen3.6-27B, reading live screenshots and emitting pyautogui-style tool calls. The paper (arXiv 2608.15930) reports 77.0 on OSWorld-Verified and 66.2 on WindowsAgentArena, beating its base by 17.7 and 24.5 points, and int...
For about a year, "run your agent locally" meant accepting a model that couldn't reliably call a tool twice in a row. That excuse is gone. Meta Superintelligence Labs published Muse Glimmer today: a 29.6B dense causal transformer, 52 layers, 6,656 hidden dim, with a ~1.8B ViT-...
Released August 28 with 78 layers, 77 of them MoE with 256 routed plus one shared expert and top-8 routing, plus a native 10B MTP layer for speculative decoding (GitHub). The attention stack uses Gated DeepSeek Sparse Attention with IndexCache for cross-layer sparse index reus...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.