ZeroHour
arXiv cs.AI / cs.LG / cs.CLpublished ()ingested Sarah Wyer

Harm Laundering in GPT Models: Evidence That Gender Discrimination Is Transformed Rather Than Reduced Across Safety-Trained Generations

infoAI safety & securityimportance 42
AI summary · glm-5.3-flash

Study of 450,000 GPT-lineage completions finds safety training transforms gender discrimination into subtle 'harm laundering' that toxicity classifiers miss.

The paper analyzes 450,000 gender-directed completions across 15 OpenAI GPT-lineage models from GPT-2 to GPT-5 under three demographic conditions. Sexual-violence clusters in GPT-2 women-directed output disappear by GPT-4, but men-directed completions gain positive representational territory women-directed output lacks, and a GPT-5 cluster (1,997 documents) frames breast cancer as a men's rights debate while three independent classifiers score it non-toxic. Women-directed topic diversity falls 36% at the GPT-4 alignment boundary (W/M=0.58 from 0.91), and representational harm disparity correlates with release date (rho=+0.55) while Detoxify toxicity does not. The authors formalize a three-criteria harm-laundering test and a three-stage detection protocol, arguing toxicity score reduction is an insufficient proxy for harm reduction.

  • Analyzes 450,000 completions across 15 GPT-lineage models
  • Toxicity scores fall as representational harm disparity grows
  • Women-directed topic diversity drops 36% at GPT-4 alignment boundary
  • GPT-5 cluster frames breast cancer as a men's rights debate
  • Proposes three-criteria test and detection protocol for harm laundering
VendorsOpenAI
OrganizationsOpenAI
Full article219 words · extracted from arxiv.org · click to collapse

Safety evaluations for large language models rely on surface-form classifiers that report declining harm scores across model generations. We provide evidence that this methodology is systematically incomplete: explicit discriminatory content is transformed rather than removed. We call this \emph{harm laundering}. Analysing 450,000 gender-directed completions across 15 models spanning GPT-2 through to GPT-5 (OpenAI GPT lineage; three demographic conditions), we show that sexual violence clusters prevalent in GPT-2 women-directed output disappear by GPT-4, while men-directed completions gain positive representational territory (caregiving, emotional range, ally identity) that women-directed completions do not. The pattern is most visible at GPT-5: Topic~5 (1,997~documents) frames breast cancer as a men's rights debate, while zero equivalent clusters appear in women-directed output. Three independent classifiers score this content as non-toxic. Sentiment scores invert at GPT-4: early models demean women; later models over-correct. Topic diversity in women-directed completions falls 36\% relative to men at the GPT-4 alignment boundary (W/M~$= 0.58$, from $0.91$ at GPT-2). REGARD representational harm disparity correlates with release date ($ρ= +0.55$, $p = .034$) while Detoxify does not ($ρ= -0.23$, $p = .42$): toxicity scores fall as representational harm grows. We formalise harm laundering as a three-criteria test and provide a three-stage detection protocol applicable to any generative model. Within the OpenAI GPT lineage, toxicity score reduction is not a sufficient proxy for harm reduction.

Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.20779