Fetching from the wire…
Public story · 2026-08-03 · high
Moonshot's guidance calls for 64-plus accelerators, but independent code now runs the weights on ordinary laptops instead.
Why now: Deltafin logged its fourth benchmark point on August 3, within a week of Moonshot's July 26 release of K3's weights.
Three independent projects put the full 2.78-trillion-parameter Kimi K3 model on single machines within a week of Moonshot's July 26 weight release.
Moonshot's own guidance calls for 64 or more accelerators to run K3. That's data-center hardware, not a desk.
Getting the same weights onto a 64GB MacBook Pro or a single M1 Max changes who can open the model up at all. Daily use is still out of reach for any of them.
WASTE, a GitHub project with 1,366 stars, keeps a 27.28GB trunk resident in memory. It streams the rest of K3's experts from NVMe using 3-bit residual vector quantization. That shrinks the model's footprint from a published 1.42TB down to 982GB on disk. It runs at 0.45 to 0.62 tokens per second.
Its own README logs a strange result. Growing the expert cache from 17.32GB to 29.32GB raised the cache hit rate from 36.2% to 41.3%. Throughput fell eightfold anyway, because a cache hit turned into a page fault.
Deltafin, a Rust project with 641 stars, published the clearest paper trail. It logged 0.0141 tokens per second on July 27, climbing to 0.2847 by August 3, a 20x jump in six days on one M1 Max.
kimi-k3-in-c goes the other direction: 176KB of plain C99, no BLAS, no framework, no GPU. It runs at 8.24GB peak RSS and 32.69 seconds per token.
These speeds aren't usable for anything real. They're proof the model runs at all, not a product.
WASTE's own numbers make that case better than any argument could. A bigger cache made the model slower, not faster, because it stopped hitting RAM and started hitting the page fault handler. Watch whether the next multi-trillion-parameter model gets a laptop port this fast, or whether K3's tricks turn out to be a one-off.
Each link below shares sources, entities, or timing with this story.
An open-weight Chinese frontier model is now a dropdown option in Microsoft's coding product. That happened before anyone finished characterizing what the model does. GitHub's changelog dated August 6 makes Kimi K3 generally available across Copilot Pro, Pro+, Max, Business an...
sqliteai/waste (769 stars, created July 28) keeps only K3's 27.28GB trunk resident and streams the ~4% of experts activated per token off SSD, opening the undistilled model in 29.06GB of RAM. The counterintuitive result is the cache table: raising the expert cache from 17.32GB...
Alongside WASTE, gavamedia/deltafin (603 stars, created July 28) runs full K3 on a single device with an OpenAI-compatible server. Moonshot published K3's open weights July 26-27 at 2.8T parameters; within 48 hours two separate projects appeared whose entire purpose is fitting...
The file is 1.56TB. That's the first thing you notice about the moonshotai/Kimi-K3 Hugging Face repo that went live today. 2.8 trillion total parameters, 104B activated, 896 routed experts with 16 selected plus 2 shared per token, 93 layers split 69 Kimi Delta Attention and 24...
Moonshot released it July 16: 2.8T parameters, MoE routing 896 experts with 16 active per token, native multimodal input, 1M-token context, with full open weights promised by July 27. Model overview here. It ships MXFP4 (4-bit float with per-block scaling) from day one, puttin...
Moonshot AI released Kimi K3, a sparse mixture-of-experts activating 16 of 896 experts per token. That's about 1.8% of the pool live at any moment, with a 1M-token context window and native vision. Two new architectural pieces show up: Kimi Delta Attention and Attention Residu...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.