Agents
EVOHARNESSBENCH moves the non-stationarity out of the task stream and into the harness, with 17 streams over 520 tools, 42 skills and 62 agents
Continual-learning benchmarks for agents normally vary the tasks while holding tools, skills and sub-agents fixed, which is backwards relative to how production systems actually change. EVOHARNESSBENCH varies the externally supplied harness along three axes (tools, skills, agents) using 17 multi-stage streams deterministically constructed from verifier-based benchmarks, totaling 802 tasks, 520 tools, 42 skills and 62 agents. It scores two things separately: whether an agent retains previously accessible competence as the harness expands, and whether accumulated experience stays useful once new capabilities appear.
Source
↳ Follow the thread