Skills
Learn a cheap rubric from a few oracle rollouts and evolve skills against it, cutting token cost 40-70%
Existing skill self-evolution methods revise skill text from execution feedback, but each oracle evaluation needs a full agent rollout, which confines search to patching whatever just failed. SkillLift treats ranking as a smoother supervision target than absolute score regression and solves a bilevel problem: an inner loop revises skills against a frozen learned rubric at no oracle cost, an outer loop spends a small number of real rollouts to re-align the rubric by rank correlation. On complex agent benchmarks it beats existing auto-skill methods with 40-70% less token cost than frontier evolving methods.
↳ Follow the thread