Fetching from the wire…
Infra2026-08-20 · source-backed
Inco AI kept parallel block drafting and added a bilinear-attention path selector scoring adjacent token pairs across the top-16 candidates for 0.6% latency overhead, plus a two-tap dynamic depthwise convolution fixing suffix decay for 3% more parameters. Result: 21% gain in mean acceptance length over DFlash, 3.1-4.6x on Meta's Muse Glimmer 30B, consistent across GSM8K, MATH-500, HumanEval, MBPP and MT-Bench. (Inco AI) The llama.cpp integration PR is open as ggml-org/llama.cpp #27342, filed 2026-08-18 and still unmerged, with community reports of up to 30% faster inference that the maintainers haven't verified.
Each link below shares sources, entities, or timing with this story.
For about a year, "run your agent locally" meant accepting a model that couldn't reliably call a tool twice in a row. That excuse is gone. Meta Superintelligence Labs published Muse Glimmer today: a 29.6B dense causal transformer, 52 layers, 6,656 hidden dim, with a ~1.8B ViT-...
Three moves, two days, no coordination between them. August 10–11: GitHub shipped Ollama as a BYOK provider inside Copilot for JetBrains (GitHub Changelog). Unsloth released Unsloth Desktop with a command literally named unsloth start claude, which points Claude Code and Codex...
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's...
The August 16 release brings video generation and image-to-video through Minimax H3, full Qwen 3.8 and Muse Glimmer support with jinja templates and tool calling, and DSpark/Dflash speculative decoding (GitHub). Practical limits rose too: max images and audio attachments to 64...
Somebody diffed the configs. Zero architectural changes. Same 64 layers, same 5,120 hidden dimension, same hybrid Gated DeltaNet → FFN / Gated Attention → FFN block structure as Qwen3.6-27B. The r/LocalLLaMA post showing this hit 945 upvotes and 157 comments, and Hugging Face...
The essay argues closed frontier labs risk concentrating too much power in too few hands, framing American open weights as the answer to Chinese open-source models (FT, corroborated by CNBC and Fortune). The sharpest lines: "I do not understand why anyone who believes that AI...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.