Agents
VLoc Bench: best of 27 models scores 0.229 File F1 on repository-scale vulnerability localization, and 38.4% of tasks get no correct answer from anything
A September 14 arXiv benchmark gives agents only a CWE description and terminal access and asks them to find the implementing files, across 500 real vulnerabilities from 290 repositories, six package ecosystems and 147 CWE categories. Across 27 language models and four static-analysis tools on a standardized agent interface, the strongest system reached File F1 of 0.229 and 38.4% of tasks received no correct localization from any evaluated system. The authors also found systems that identify vulnerabilities well still report unsupported locations on already-patched repos, separating localization from detection as a distinct capability.
Source
↳ Follow the thread