Skills
FORGE treats each failure as two signals and synthesizes new training data alongside the prompt revision, for 16.52 points
Automatic prompt optimization normally holds training data fixed, so repeated optimization only sees weaknesses already present in those instances. FORGE abstracts imperfect executions into reusable failure modes and synthesizes new training data through four mutation strategies, feeding verified instances back into prompt search so updated prompts expose the next data need. Across eight benchmarks it improves the aggregate score 16.52 points over the unoptimized baseline, and the synthesized data transfer, improving all nine APO comparisons by 2-9 points and all three GRPO comparisons by 4-8 points under matched budgets.
↳ Follow the thread