Sources
ASPIRE benchmarks self-evolution from a vague goal with the evaluation tasks hidden from the agent
arXiv 2608.31111 (2026-08-31, 27 HF upvotes) argues existing self-evolution work is not really self-directed, because humans hand the agent tasks and metrics, reducing the problem to optimizing a stated objective. ASPIRE supplies only a natural-language capability goal such as 'become a better physicist' while the downstream evaluation stays hidden, forcing the agent to pick data and update methods, construct its own training and validation signals, and decide when to stop. It supports both model-weight and agent-harness evolution in one interactive environment and grades on a hidden expert-authored set of 520 items.
Source
↳ Follow the thread