Research
Deleting the Boilerplate Refusal Sentence From Safety-Tuning Data Cuts False Refusals Without Losing Safety
arXiv 2609.04714 decomposes safety-tuning responses into a boilerplate refusal statement and a rationale explaining the refusal, then tests which component drives behavior. The refusal statements turn out to impede discrimination between genuinely harmful queries and benign ones with superficially risky wording, by inducing reliance on surface cues, which is why models fumble the difference between shooting someone and shooting a photo. Training on rationales alone reduced false refusals while keeping safety performance comparable, and the effect also appeared in an in-context-learning configuration.
↳ Follow the thread