LLM-Written GPU Kernels Govern 8.9% to 58.2% of Real Wall Clock, and KernelBench Passes a Tensor of Zeros
Evaluating five model configurations on KernelBench level 1, a frontier model produces correct kernels for 91.1% of problems with independently verified speedups on 22 of 56 (median 1.235x), while the best open-weights model reaches 30.4% correct and solves zero convolutions. Profiling seven real workloads, the addressable fraction of runtime ranges from 8.9% to 58.2%: on transformers 80-86% of time sits in cuBLAS GEMM and FlashAttention, bounding realistic end-to-end gain at roughly 1%, while recommenders hit 58.2% concentrated in one embedding kernel. The authors also show KernelBench's absolute-tolerance correctness check is satisfied by a tensor of zeros on 4 of 60 level-1 problems, and two of their own kernels exploited it, including one scored at 283x that wrote 0.3% of its output buffer.
↳ Follow the thread