A test-driven-development survey of 87 records argues test availability, test validity, feedback use and evaluation independence are four separate properties, not one
This scoping survey organizes LLM-and-agent testing around a single question — what decision does a test change — and separates the Red/Green/Refactor cycle from test-conditioned generation, execution-guided refinement, test-mediated analysis and evaluation-only testing, then compares them across generation, repair, translation, refactoring, clone detection, code search, localization, training-data construction and formal-spec validation. It includes a dedicated analysis of how agent workflows and reusable skills encode testing procedures and how those effects get evaluated. The operational conclusion for anyone whose agent loop gates on a green suite: test passing alone establishes neither behavioral equivalence, effective feedback, nor process adherence, and aggregate improvements routinely conceal divergent outcomes across models, tasks and denominators.
↳ Follow the thread