Fetching from the wire…
Infra2026-08-27 · source-backed
The August 26 release adds Decode Context Parallel for Kimi-K3, fused FlashKDA decode and prefill kernels, combined all-gathers with a claimed 1.5-3x kernel-level speedup, an adaptive speculative token budget worth about 60% better DSpark TTFT, and optional shared-expert sharding saving about 17 GiB per GPU (GitHub). DeepSeek V4 sparse MLA now works end-to-end across plain decode, MTP and DSpark speculative decoding, with AMD Quark NVFP4 and ROCm enablement on gfx11 and gfx950. If you self-host frontier open weights, this is the release that makes the two newest Chinese flagships cheap to serve.
Each link below shares sources, entities, or timing with this story.
The file is 1.56TB. That's the first thing you notice about the moonshotai/Kimi-K3 Hugging Face repo that went live today. 2.8 trillion total parameters, 104B activated, 896 routed experts with 16 selected plus 2 shared per token, 93 layers split 69 Kimi Delta Attention and 24...
A Chinese lab shipped a runtime that manages two American coding agents as subagents, and it went from repo creation to 145,439 stars in four days. deepseek-ai/deepseek-harness published dsh-v0.1.0-rc.7 at 12:01 UTC today, its first tagged release since the repo appeared on Au...
The assumption that proprietary models own the coding benchmark crown just broke. Moonshot AI's Kimi K2.6 leads on 5 of 8 major agentic coding benchmarks while being the only open-weight model in the top tier. SWE-Bench Pro: 58.6% vs GPT-5.4's 57.7% and Claude Opus 4.6's 53.4%...
DeepSeek-V4-Flash-0731 landed July 31 under MIT with a DSpark speculative-decoding module attached. Terminal Bench 2.1: 82.7. Toolathlon-Verified: 70.3. DSBench-FullStack: 68.7. DeepSWE: 54.4. NL2Repo: 54.2. The model card claims it beats DeepSeek-V4-Pro (Preview) "despite its...
OpenAI cut GPT-5.6 Luna roughly 80%, from $1 to $0.20 per million input and $6 to $1.20 output. Anthropic priced Opus 5 at $5/$25 per million, half of Fable 5. The trigger is DeepSeek, Zhipu's GLM-5.2 and Moonshot's Kimi K3 landing 60-90% below US flagship pricing, with DoorDa...
DeepSeek released DSpark plus a full open-source training and eval stack called DeepSpec. It adds a lightweight serial module atop a parallel draft model to fix acceptance-rate decay, lifting accepted-token length 16.3% to 30.9% over Eagle3 and DFlash. The paper hit #1 on HN a...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.