Beyond Isolated Tasks: Framework for Evaluating Coding Agents on Sequential Software Evolution
arXiv·medium signal
Existing coding agent benchmarks evaluate performance on isolated, single-PR tasks in a stateless manner, failing to capture real-world software development where code changes accumulate, technical debt accrues, and tasks depend on prior context. This paper introduces a framework for evaluating agents on sequential software evolution — multi-step, stateful development workflows that better reflect how engineers actually work with codebases over time.