A 384-Point Teardown Measures Why Local LLMs Degrade: NVFP4 Hits ~50% Top-1 Token Disagreement by 88k Context
Level1Techs Forum / Hacker News·high signal
A Level1Techs writeup that hit 384 points and 144 comments on Hacker News captured full-vocabulary logits and computed KL divergence in FP64 to trace where local inference diverges from the reference model. Testing Qwen 3.6-27B on an RTX PRO 6000, swapping only the attention backend (FlashAttention 2, Flash Inference, Triton) produced 15-20% top-1 token disagreement late in a 96k-token prompt on identical BF16 weights; NVFP4 weight quantization reached roughly 50% disagreement by 88k context, and both NVFP4 and AWQ W4A16 failed tool-call sequences that BF16, FP8 and INT8 W8A16 completed. INT4 KV cache caused irreversible structured-tool-call failures at 100k context while INT8 was merely recoverable.