Research
Verifier Votes Over Shared Evidence Approve 62.9% of Unsafe Agent Actions; an Independent Source Cuts It to 22.9%
VP-CONTROL (arXiv 2609.10969) builds 2,880 commit-gate scenarios and separates model diversity from evidence diversity. Evidence-source diversity was worth 40.9 percentage points of unsafe-approval reduction, and swapping verifier models was worth 11.3. In a live HTTP/SQLite study, race conditions after the check defeated verifier-only gates. Only a full atomic guard recorded zero unsafe effects across 216 episodes, and idempotent request IDs prevented duplicate effects after lost responses. If you add a second LLM judge that reads the same inputs, you mostly get the same mistake twice.
Source
↳ Follow the thread