← The Wire
Source trail

ByteShape (via r/LocalLLaMA, 126 upvotes)

Public MindPattern findings, entities, and graph evidence that cite this source.

Findings
1
All-time hits
1
High value
0
Last seen
2026-09-16

Related findings

  1. 2026-09-16 / REDDITByteShape's Qwen3.8-27B quants hit 99.63% of BF16 at 3.84 bpw, and argue KL divergence is the wrong quant metricByteShape released full ShapeLearn GGUFs for Qwen3.8-27B (repo last modified September 15, 17,232 downloads, 60 likes) claiming 3.84 bpw reaches 99.63% of BF16's aggregate score across eight benchmarks and 3.23 bpw reaches 98.72%, with all five models on the measured quality/speed-bpw frontier across six GPUs against Unsloth v3, ISTA-DASLab, AtomicChat and Bartowski. The load-bearing argument is methodological: Unsloth Dynamic V3's UD-IQ3_S posted ~20% lower KLD at comparable size yet lost on downstream task results, so builders picking quants by KLD alone may be picking the wrong one. They also report DFlash2 at 1.34-2.10x baseline throughput and MTP at 1.28-1.66x under temperature sampling.
Open latest cited source