Pre-registered ablation shows removing an LLM verifier stage from an offensive-security agent shifts median reported findings from 0 to 2 per run
A September 14 arXiv paper runs a 15-run pilot, a pre-registered 20-run confirmatory ablation and a pre-registered 2x2 factorial with 40 runs across two vulnerable lab systems, testing whether a verifier-and-acceptance stage changes what an LLM offensive-security agent reports. Removing verification raised reported findings (median 2 versus 0 per run, p = 0.00003) and cut precision (0.353 versus 0.471, p = 0.0087), with the model verifier itself rather than the deterministic rules accounting for the suppression (p = 0.004) and recall unchanged (p = 0.158). The full design retained 93.8% of model-adjudicated true candidates but missed its pre-registered non-inferiority bound, and human verification is still pending.
Source
↳ Follow the thread