Evaluating LLM-Based Test Generation Under Software Evolution: Generated Tests Break Across Versions
arXiv·medium signal
Researchers evaluate whether LLM-generated unit tests remain valid as software evolves across versions. The study reveals that tests generated by current LLMs frequently break under software evolution — a critical finding for coding agents that generate tests as part of CI/CD workflows, since tests must survive real-world code churn to provide lasting value.