Reddit
Qwen3.8-27B Stays Usable Down to Q3_XXS, Fitting a 16GB RTX 4060 Ti at 30-35 Tokens/Sec
A r/LocalLLaMA user who normally refuses Q3 quants ran Qwen3.8-27B at Q3_XXS on a 16GB RTX 4060 Ti and reports one-shot successes on coding tasks that Qwen 3.6 35B either failed outright or needed days of prompting to finish. Fully in VRAM it runs 30-35 tok/s, matching the speed of a higher quant of the 35B offloaded to system RAM. The thread has 113 comments arguing the point, so treat it as a strong signal rather than a settled result, but the practical read is that the 27B's quant floor is lower than the 35B's was.
↳ Follow the thread