Research
Five of Nine Lightweight Guardrail Models Flip Malicious to Benign Just by Repeating the Prompt
Overflip is a repetition-induced instability in compact guardrail classifiers (DeBERTa-class backbones trained at 512 tokens with bucketed relative positional encodings). On a 100-prompt benchmark, five of nine widely used guardrails flip MAL to BEN as the input lengthens, with flip rates from 8% to 92% and first flips at roughly 2.6k to 9.4k tokens. The malicious content is preserved intact; repetition homogenizes token-level attention over repeated structure, a different trajectory from classic attention-dilution padding attacks.
↳ Follow the thread