smolbenchmark ranks sub-8GB models by tokens per joule and thermals on hardware you already own
yuvrajsingh-mist.github.io (via r/LocalLLaMA)·medium signal
Released to r/LocalLLaMA today, smolbenchmark inverts the usual leaderboard: 13 model families that fit in 8GB, ranked by decode speed, tokens per joule and heat, measured on tablets, phones, Macs, Jetsons and Raspberry Pis rather than a server. About 1,000 configs are live for the Jetson Orin Nano Super 8GB, capturing tok/s, tok/J, inter-token latency, power, thermals and battery, with Pi, phone and Mac mini runs still pending. The top comment is the useful caveat: without a reproducibility block (quant, backend version, context length, warmup count, plugged-in state) a cold first run flatters a model and a long run reverses the ranking through throttling.