Reddit
ISTA-DASLab's GSQ-RCO quants shrink Qwen3.8-Flash-Next to 68-76GB, with a Q2_0 built for throughput over accuracy
ISTA-DASLab shipped GSQ-RCO GGUFs for Qwen3.8-Flash-Next (HF repo last modified September 15, 4,829 downloads), cutting the model from roughly 80-95GB to 68-76GB. Their IQ3_XXS operating point matches the base model exactly on AIME25 (100.00) and is within 0.51 on GPQA-Diamond and 1.14 on LiveCodeBench v6 at about a fifth of BF16 size. The unusual variant is Q2_0, which deliberately avoids lookup-table quant formats so decode cost stays flat, buying 3.4x prompt throughput and 1.9x lower end-to-end latency versus IQ2_XS for 0.09 points of task average. One commenter on dual 3090s reports it is currently slower in practice for them because tensor parallel and MTP support are missing.
↳ Follow the thread