Reddit
A solo WebGPU engine doubled 1-bit 27B decode speed in two days to 30 tok/s on a 6 GB laptop GPU
mentria.ai, a browser inference engine written from scratch in WebGPU/WGSL, is running Prism ML's natively 1-bit Bonsai-27B at 25 to 30 tok/s on an RTX 3060 Laptop with 6 GB of VRAM, in Chrome, with nothing installed and nothing leaving the machine. At one sign bit per weight and one FP16 scale per 128 weights, roughly 1.14 bits per parameter, the 27B fits in 3.8 GB. Two days before the post the same model decoded at 15 tok/s on the same laptop; the author reports 804 GPU dispatches per token, 401 of them the 1-bit matvec kernel, and published his own eval deltas (GSM8K 90.5, MMLU-Redux 82.9, MATH-500 98.1) rather than quoting Prism's.
↳ Follow the thread