Noise Out, Bias In: Targeted Bias Injection in Diffusion Language Models via Closed-Loop Activation Steering
A closed-loop steering attack injects targeted demographic bias into diffusion language models during denoising.
Masked diffusion language models re-expose token distributions at every denoising step before a token is committed. The authors use a proportional-integral controller to adapt an activation steering vector toward an adversary-chosen demographic answer. On ambiguous BBQ questions, LLaDA-8B-Instruct's preference for the targeted group rose from 1.8 to 16.7 percentage points, and on SocialStigmaQA stigmatizing answers rose from 17.6% to 58.1%. Constant-strength steering shifted answers less and corrupted nearly three times as many outputs; each attack took about 40 minutes on one GPU.