Skill optimization via contextual bandits cut optimization cost 55-58% using only 50 examples per benchmark
COBRA-Skills treats agent skill optimization as budgeted sequential optimization over a candidate space that keeps evolving, pairing contextual-bandit-guided prioritization with evidence-grounded skill evolution so evaluation budget goes to promising or informative candidates rather than the whole population. Across six agent benchmarks and three target models it achieved the strongest average performance among compared methods while reducing optimization cost 55-58% against SkillOpt, on 50 unique optimization examples per benchmark. It stayed robust when the agent harness changed and worked when the target model itself generated and refined the skills, which removes the need for a separate stronger teacher.
↳ Follow the thread