Fetching from the wire…
Infra2026-08-02 · source-backed
Wafer.ai ran K3 at TP8 on a single MI355X node: 952 tok/s aggregate, 118 tok/s single-stream decode, ~13k tok/s steady-state prefill after fixing missing PyTorch sampling functions and AITER MLA prefill kernels via SGLang and ROCm. Against a two-node TP16 B200 deployment that's 3.8x aggregate throughput per node. B300 delivers 1.65x higher aggregate but costs 2.4x more per GPU. Vendor-adjacent blog, not an independent benchmark house: weight accordingly.
Each link below shares sources, entities, or timing with this story.
A practitioner ran Kimi K3 on 8x B300 via Modal at $56.79/hour with vLLM, TP8 and native MXFP4: 27-minute cold boot for a 1.56 TB load, TTFT 0.92 to 1.02s, 92 tok/s steady decode, roughly $36 of GPU time per clean run and $1,363/day left warm. The cheaper path was worse. Unslo...
The UK AI Security Institute and US CAISI published a joint preliminary cyber evaluation of Moonshot's Kimi K3 (open weights due July 27). On ExploitBench, a Carnegie Mellon benchmark covering 41 post-2023 Chrome V8 vulnerabilities, K3 hit 32% versus 76% for the most cyber-cap...
Moonshot released K3's open weights July 26 with official guidance calling for 64+ accelerators. WASTE (1,366 stars, created July 28) runs it on a 64GB MacBook Pro at 0.45-0.62 tok/s, keeping the 27.28GB trunk resident and streaming experts from NVMe with 3-bit residual vector...
Moonshot exposes an Anthropic-compatible endpoint, so pointing Claude Code at K3 means setting the Anthropic base URL and supplying a Moonshot key. No new CLI, no config rewrite. Hosted at $3/$15 per Mtok, same tier as Claude Sonnet 4.6, and Artificial Analysis scores K3 at 57...
The file is 1.56TB. That's the first thing you notice about the moonshotai/Kimi-K3 Hugging Face repo that went live today. 2.8 trillion total parameters, 104B activated, 896 routed experts with 16 selected plus 2 shared per token, 93 layers split 69 Kimi Delta Attention and 24...
sqliteai/waste (769 stars, created July 28) keeps only K3's 27.28GB trunk resident and streams the ~4% of experts activated per token off SSD, opening the undistilled model in 29.06GB of RAM. The counterintuitive result is the cache table: raising the expert cache from 17.32GB...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.