Vibe Coding
Tip: Intel Optane Persistent Memory Enables 1 Trillion Parameter Model Locally at 4 tok/s on a Single Machine
A builder on r/LocalLLaMA demonstrated running Kimi K2.5 (1 trillion parameters) locally at ~4 tokens/second using Intel Optane Persistent Memory. The key insight is that Optane DCPMM provides massive, byte-addressable memory capacity at a fraction of DRAM cost, enabling model sizes previously impossible outside data centers. With 605 upvotes and 99 comments, the post includes full build specs. For anyone running large local models for coding, this is the most cost-effective path to running frontier-scale models without cloud API dependency.
Source
↳ Follow the thread