Reddit
Self-Hosting Kimi K3's 2.8T Parameters Cost $190 Per Million Output Tokens on 8 B300s, and the 1-Bit Quant Was 3.3x Worse
A practitioner ran Kimi K3 on 8x B300 via Modal at $56.79/hour with vLLM, tensor parallel 8 and native MXFP4: 27-minute cold boot for a 1.56 TB load, TTFT 0.92 to 1.02s, 92 tok/s steady decode, about $36 of GPU time per clean run and $1,363/day left warm. The cheaper path was worse, not better. Unsloth's 1-bit UD-IQ1_S at 594 GB on 8x A100-80GB via llama.cpp cost $19.99/hour but delivered roughly 9 tok/s and about $620 per million tokens. Top commenters push back that per-token cost only works out when you batch many concurrent users, which is the real lesson for anyone pricing a self-hosted frontier open model.
Source
↳ Follow the thread