Fetching from the wire…
Public story · 2026-07-27 · high
Any service provider clearing $20M in revenue over any 12 months must sign a separate license with Moonshot.
Why now: Kimi K3's weights, license terms, and day-0 vLLM support all landed on July 27, and every GGUF quant published since fails to load in llama.cpp.
Moonshot released Kimi K3's weights on July 27, a 2.8-trillion-parameter model that drops the modified-MIT license its predecessor K2 shipped under, per the model card.
Any company running K3 as a paid service must sign a separate agreement with Moonshot. That requirement kicks in once revenue crosses $20 million over any 12-month span, per the model card. That number should shape a product team's build decision, not the trillion-parameter headline.
The model activates 104 billion of its 2.8 trillion parameters per token across 896 routed experts. It runs 93 layers split between Kimi Delta Attention and Gated MLA, with a 1,048,576-token context window. It shipped as MXFP4 weights trained quantization-aware, not squeezed down after the fact, per Willison's rundown of the Hugging Face repo.
The license is labeled 'open weight,' not open source, and Willison flagged three secondary write-ups mislabeling K3 as Apache 2.0, calling them wrong. K2's requirement to credit Moonshot in consumer products is gone in K3.
An r/LocalLLaMA deployment team worked the hardware math first. 8x A100 80GB gives 640GB against 1.4TB of weights, three nodes before a byte of KV cache. Ampere has no FP4 or FP8 tensor cores. 8x H200 hits about 1.13TB, still two nodes. Only 8x B300, at 2.3TB, fits the model on one node, the hardware Moonshot quantized for.
vLLM shipped day-0 support in v0.26.0, hitting 118 tokens per second per user baseline. DSpark speculative decoding pushed that to 370 tokens per second on 16x GB300 NVL72, a 3.14x jump. It runs only in Docker, since the build needs a pre-release FlashInfer.
No GGUF quant of K3 loads on llama.cpp. The thread tracking its support lists a missing architecture registration, an unimplemented activation function, and unhandled tensors for its latent expert space. Builds from GrEarl, Kuberwastaken, and AtomicChat all fail to load.
Moonshot's tech report claims Quantile Balancing and Per-Head Muon tuning give K3 2.5 times better compute-to-capability than K2. The hfviewer Expert Atlas post found K3's active parameter fraction is 1.8%, the smallest yet on the largest expert pool shipped in the open.
Each link below shares sources, entities, or timing with this story.
DeepSeek released V4 on April 24 and the numbers demand attention. V4-Pro is 1.6 trillion parameters total with 49 billion active, MIT-licensed, native 1M-token context. It scores 80.6% on SWE-bench Verified, putting it within 0.2 points of Claude Opus 4.6. On Terminal-Bench 2...
DeepSeek released V4 Preview on April 24 with two open-weight variants: V4-Pro (1.6T total parameters, 49B activated via MoE) and V4-Flash (284B parameters, 13B activated). Both support 1M-token context windows. Both are Apache 2.0 licensed. Both are live right now on Hugging...
A practitioner ran Kimi K3 on 8x B300 via Modal at $56.79/hour with vLLM, TP8 and native MXFP4: 27-minute cold boot for a 1.56 TB load, TTFT 0.92 to 1.02s, 92 tok/s steady decode, roughly $36 of GPU time per clean run and $1,363/day left warm. The cheaper path was worse. Unslo...
For about a year, "run your agent locally" meant accepting a model that couldn't reliably call a tool twice in a row. That excuse is gone. Meta Superintelligence Labs published Muse Glimmer today: a 29.6B dense causal transformer, 52 layers, 6,656 hidden dim, with a ~1.8B ViT-...
Cursor shipped Composer 2 on March 19, marketing it as a proprietary in-house model. Within 24 hours, a developer found the API routing to kimi-k2p5-rl-0317-s515-fast. The model powering the most-hyped coding tool update of the month was Moonshot AI's Kimi K2.5 with continued...
PR #26062, "server: support MCP stdio," by ngxson, merged into ggml-org/llama.cpp on July 25 (r/LocalLLaMA). It landed alongside #26061 (vendored subprocess.h, merged July 24) and pwilkin's #26075 integration-and-tests PR. Until now, llama-server's web UI could only talk to MC...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.