Fetching from the wire…
Models2026-08-29 · source-backed
Two methods: GSQ (Gumbel-Softmax Quantization), post-training scalar quantization that jointly learns grid assignments and scales at 2-3 bits, and RCO (Riemannian Constrained Optimization), assigning a quant type per tensor under a strict size budget by gradient descent on the task loss. The 3.00 bpw build (10.1 GB) matches the BF16 base on AIME25 at 100.00, comes within about a point on GPQA-Diamond (88.89 against 89.90) and LiveCodeBench v6 (84.57 against 85.71). The 2.75 bpw build (9.3 GB) still reaches AIME25 100.00 with a zero-shot average exceeding BF16. All three files run unmodified in llama.cpp, Ollama and LM Studio. Given today's GGUF filename audit, verify the bpw yourself before quoting these. (r/LocalLLaMA)
Each link below shares sources, entities, or timing with this story.
For about a year, "run your agent locally" meant accepting a model that couldn't reliably call a tool twice in a row. That excuse is gone. Meta Superintelligence Labs published Muse Glimmer today: a 29.6B dense causal transformer, 52 layers, 6,656 hidden dim, with a ~1.8B ViT-...
PR #26062, "server: support MCP stdio," by ngxson, merged into ggml-org/llama.cpp on July 25 (r/LocalLLaMA). It landed alongside #26061 (vendored subprocess.h, merged July 24) and pwilkin's #26075 integration-and-tests PR. Until now, llama-server's web UI could only talk to MC...
The models are good. The license is the real story. Google released Gemma 4 on April 2 with four variants: E2B, E4B, 26B MoE, and 31B Dense. All built on the Gemini 3 architecture. The 31B Dense variant claimed #3 on Arena AI's text leaderboard, beating models 20x its size. Th...
Alibaba released Qwen3.6-27B on April 22. Dense architecture. Open weights. 77.2% on SWE-bench Verified, within 3.7 points of Claude Opus 4.6. On SkillsBench, it scores 48.2% versus its own 397B MoE predecessor's 30.0%. That's a 77% improvement with 14.8x fewer parameters. Let...
The repo appeared August 24, opening the weights of a multimodal MoE that Qwen frames explicitly as an architecture preview, the same role Qwen3-Next played for Qwen3.5 (GitHub). The hybrid Gated DeltaNet plus Gated Attention design it previews already carried through the Qwen...
Hugging Face published its Summer 2026 State of Open Models report on August 14, and one statistic in it went almost entirely unremarked in the coverage. By July 2026, agents rather than humans became the Hub's primary users. Claude Code alone accounted for 44.4% of all agent...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.