FrogNano trains a competitive 4B coding agent on synthetic tasks alone, with no distillation from a larger model
arXiv·medium signal
arXiv 2609.07925 (submitted 2026-09-07) post-trains a 4B model purely by RL across roughly 1,500 SWE environments, using an online task synthesis pipeline that generates tasks calibrated to the current checkpoint's learnability frontier. The claim is that the frontier calibration, not the data volume, is what makes synthetic-only training work. For builders the interesting consequence is a coding agent small enough to run on minimal hardware without a teacher model in the loop.