ZeroHour
AI model

GPT-2

1 mentions in 7 days · 2 in 30 days · 2 total · first seen · last

Timeline

Harm Laundering in GPT Models: Evidence That Gender Discrimination Is Transformed Rather Than Reduced Across Safety-Trained Generations

Study of 450,000 GPT-lineage completions finds safety training transforms gender discrimination into subtle 'harm laundering' that toxicity classifiers miss.

The paper analyzes 450,000 gender-directed completions across 15 OpenAI GPT-lineage models from GPT-2 to GPT-5 under three demographic conditions. Sexual-violence clusters in GPT-2 women-directed output disappear by GPT-4, but men-directed completions gain positive representational territory women-directed output lacks, and a GPT-5 cluster (1,997 documents) frames breast cancer as a men's rights debate while three independent classifiers score it non-toxic. Women-directed topic diversity falls 36% at the GPT-4 alignment boundary (W/M=0.58 from 0.91), and representational harm disparity correlates with release date (rho=+0.55) while Detoxify toxicity does not. The authors formalize a three-criteria harm-laundering test and a three-stage detection protocol, arguing toxicity score reduction is an insufficient proxy for harm reduction.

When Does Scale-Invariant Optimization Become Unstable? An Exact Schedule Law with Weight Decay

Researchers derive an exact law linking learning-rate schedules and weight decay in normalized networks, pinpointing when scale-invariant optimization destabilizes.

The paper shows that normalization makes large parts of neural networks scale-invariant, creating a hidden feedback loop where learning-rate schedules and weight decay interact through the parameter norm to control the effective optimizer step. An exact discrete-time law with a single scalar quantity separates contraction- and expansion-dominated effective learning-rate regimes, and the balance point is intrinsically unstable, so constant learning rate with weight decay produces recurrent behavior instead of a stable equilibrium. A unified homogeneous-optimizer framework explains why adaptive methods stabilize more weakly under normalization. The law is validated with high precision on MLPs, CNNs, and GPT-2 across MNIST, CIFAR, WikiText, and OpenWebText, with code released on GitHub.

arXiv cs.AI / cs.LG / cs.CL · 9d agoAI research

Appears with

Entities are extracted by the model from each article. Watching an entity keeps it in this browser only (no account); the watchlist page and dashboard alerts use it.