Swift-Qwen3.8-27B cuts thinking tokens 58.3% with on-policy distillation and under 1% accuracy loss
UkisAI post-trained Qwen 3.8 27B by identifying tokens tied to overthinking and penalizing those specifically rather than capping reasoning length, then repaired accuracy with on-policy distillation, reporting 58.3% fewer thinking tokens, a 1.95x speedup, and under 1% accuracy loss versus xhigh effort. The Hugging Face API shows the repo was created 2026-09-08 and last modified 2026-09-13, with 1,355 downloads and 198 likes. Worth checking before you adopt it: the license is a custom 'Swift Open License 1.0', not Apache or MIT, despite the post describing the release as open-sourced. There is a free 5 RPM OpenAI-compatible research API on Nvidia-donated GPUs, plus official Q1 to Q8 GGUFs and Bartowski quants.
Source
↳ Follow the thread