Research
A Prompt-Injection Detector Scoring F1 0.98 In-Distribution Misclassifies a Third of Security-Adjacent Benign Prompts
PIDS-Bench evaluates seven detectors at fixed thresholds across in-distribution inputs, hard-benign prompts that mimic injection structure, obfuscated attacks, and domain and structural shifts, scoring false positives as a first-class axis rather than folding them into aggregate F1. A detector exceeding F1 = 0.98 on its held-out split still misclassifies roughly one-third of an externally-sourced benign subset restricted to security-adjacent content. Across a full threshold sweep and five training seeds, no internal detector reached an operating point with F1 >= 0.95 and hard-benign FPR <= 0.10 at the same time.
↳ Follow the thread