Reddit
Kernel fusion took GLM 5.3 Flash from 29 to 40 tok/s on an M3 Ultra by removing idle GPU gaps, not by better matmuls
The ds4 maintainer found GLM-5.3-Flash was single-stream decoding at only 59% of the M3 Ultra's measured memory bandwidth, and the bottleneck was not the big weight-streaming kernels but dozens of small kernels between them each paying dispatch latency. Fusing that work into larger dispatches produced 29 to 40 t/s at short context and 24 to 38 t/s at 62k, hitting about 81% of the measured bandwidth ceiling, with a Claude Code harness at 200k depth averaging over 38 t/s output. This is a concrete demonstration that on Apple Silicon the remaining headroom is in dispatch count, not arithmetic.
↳ Follow the thread