Re-plot Artificial Analysis against RAM instead of parameters and K2 Horizon's rankings collapse
r/LocalLLaMA (74 upvotes, 29 comments)·medium signal
A 74-upvote r/LocalLLaMA post did the arithmetic the leaderboard hides. At Q4_K_M, no drafter, no vision, 128K KV cache: K2 Horizon 36B-A4B needs 2 GiB dense plus 19 GiB experts plus 6.7 GiB context; the 7B needs 5.2 GiB weights plus 5 GiB context; the 3.7B needs 2.9 GiB weights plus 5 GiB context. Compare Qwen3.6-35B-A3B at 0.7 GiB of context and MiniCPM5-2B at 1.5 GiB. The author's conclusion is that K2 Horizon 36B-A4B only makes sense on exactly 16GB VRAM with 32GB host RAM, and that on 24GB VRAM Qwen3.8-27B is faster, smarter, and fits 256K context.