Harm Laundering in GPT Models: Evidence That Gender Discrimination Is Transformed Rather Than Reduced Across Safety-Trained Generations
Study of 450,000 GPT-lineage completions finds safety training transforms gender discrimination into subtle 'harm laundering' that toxicity classifiers miss.
The paper analyzes 450,000 gender-directed completions across 15 OpenAI GPT-lineage models from GPT-2 to GPT-5 under three demographic conditions. Sexual-violence clusters in GPT-2 women-directed output disappear by GPT-4, but men-directed completions gain positive representational territory women-directed output lacks, and a GPT-5 cluster (1,997 documents) frames breast cancer as a men's rights debate while three independent classifiers score it non-toxic. Women-directed topic diversity falls 36% at the GPT-4 alignment boundary (W/M=0.58 from 0.91), and representational harm disparity correlates with release date (rho=+0.55) while Detoxify toxicity does not. The authors formalize a three-criteria harm-laundering test and a three-stage detection protocol, arguing toxicity score reduction is an insufficient proxy for harm reduction.