Safety Scores Fell Across GPT Generations While Representational Harm Grew
arXiv 2609.20779 (17 Sep 2026) analyzes 450,000 gender-directed completions across 15 models from GPT-2 through GPT-5 and argues that surface-form classifiers report declining harm because explicit content is transformed rather than removed, which the authors call harm laundering. Sexual violence clusters common in GPT-2 women-directed output vanish by GPT-4 while men-directed completions gain positive representational territory that women-directed completions do not, and topic diversity for women falls 36% relative to men at the GPT-4 alignment boundary. REGARD representational harm disparity correlates positively with release date (rho = +0.55, p = .034) while Detoxify does not (rho = -0.23, p = .42), meaning toxicity scores fall as representational harm grows.
↳ Follow the thread