Sources
RetireOPD lets the student fire its teacher when the gap stops closing, beating the teacher every time
RetireOPD (arXiv 2609.20784, Sept 17) attacks two failures in self on-policy distillation for multi-turn agents: privileged information alone does not make a teacher reliable, and teacher supervision helps only at certain training stages. Instead of a fixed distillation schedule, the student drops the teacher on its own once their discrepancy stops shrinking and it reaches a target fraction of the teacher's success rate, then continues on RL alone. Across Qwen2.5 models from 1.5B to 7B it improves ALFWorld success by 14.1-18.8% and WebShop accuracy by 11.8-19.0% over the RL baseline, and surpasses its own skill-conditioned teacher in every setting.
Source
↳ Follow the thread