Reddit
100 zebra puzzles and 6.5 minutes on one H100 moved Qwen 3 4B Base from 54% to 84.6% on MATH-500
A Hugging Face blog post published 2026-09-11 introduces PCSS, a per-example calibrated sigmoid scaler derived from KTO that reduces to standard SFT gradient times a scaler decaying as the model masters an example. Fine-tuning Qwen 3 4B Base on 100 zebra puzzles for 6.5 minutes on a single H100/H200 gave 84.60% on MATH-500, a +30.5% delta, with +2.57% on AIME 2025. The gains shrink sharply with model strength: Granite 4.1 3B got +11.1% and Qwen 3.5 9B only +3.1%, so this is a small-model recipe, not a general one. A reproduction notebook and the dataset viewer are both public.
↳ Follow the thread