Fetching from the wire…
Infra2026-08-02 · source-backed
sqliteai/waste (769 stars, created July 28) keeps only K3's 27.28GB trunk resident and streams the ~4% of experts activated per token off SSD, opening the undistilled model in 29.06GB of RAM. The counterintuitive result is the cache table: raising the expert cache from 17.32GB to 23.32GB lifts hit rate from 36.2% to 38.4% while throughput collapses eightfold, because a cache hit becomes a page fault. Storage bandwidth is binding: a cold token reads ~17GB of experts, and a USB enclosure at 0.94 GB/s versus internal NVMe at 12.78 GB/s is the difference between working and not. Layers validate against PyTorch to 3.6e-06 on final logits.
Each link below shares sources, entities, or timing with this story.
Moonshot released K3's open weights July 26 with official guidance calling for 64+ accelerators. WASTE (1,366 stars, created July 28) runs it on a 64GB MacBook Pro at 0.45-0.62 tok/s, keeping the 27.28GB trunk resident and streaming experts from NVMe with 3-bit residual vector...
An open-weight Chinese frontier model is now a dropdown option in Microsoft's coding product. That happened before anyone finished characterizing what the model does. GitHub's changelog dated August 6 makes Kimi K3 generally available across Copilot Pro, Pro+, Max, Business an...
The file is 1.56TB. That's the first thing you notice about the moonshotai/Kimi-K3 Hugging Face repo that went live today. 2.8 trillion total parameters, 104B activated, 896 routed experts with 16 selected plus 2 shared per token, 93 layers split 69 Kimi Delta Attention and 24...
Moonshot released it July 16: 2.8T parameters, MoE routing 896 experts with 16 active per token, native multimodal input, 1M-token context, with full open weights promised by July 27. Model overview here. It ships MXFP4 (4-bit float with per-block scaling) from day one, puttin...
diegosouzapw/OmniRoute added 1,343 stars on July 20, a single MIT-licensed gateway across 268+ providers (50+ free) and 500+ models including Claude, GPT, Gemini, Kimi K3, GLM and DeepSeek, wired for Claude Code, Codex, Cursor, Cline and Copilot. Quota-aware automatic fallback...
The UK AI Security Institute and US CAISI published a joint preliminary cyber evaluation of Moonshot's Kimi K3 (open weights due July 27). On ExploitBench, a Carnegie Mellon benchmark covering 41 post-2023 Chrome V8 vulnerabilities, K3 hit 32% versus 76% for the most cyber-cap...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.