Separate LLaDA-8B papers on distillation and bias steering
October 2026 reports on LLaDA-8B-Instruct describe Pivot-SD self-distillation with 200 questions and closed-loop steering that raises targeted bias.
Two early-October 2026 reports both involve the masked diffusion model LLaDA-8B-Instruct, but they describe different methods and do not contradict each other. Pivot-SD, posted on arXiv on 2 October 2026, is an offline self-distillation framework that supervises only pivot denoising commitments selected by how much they reduce uncertainty over remaining masked positions. Pivots from successful trajectories are trained with cross-entropy and failed-trajectory pivots with unlikelihood loss; using 200 questions and four rollouts each, the authors report gains over full-sequence supervised fine-tuning and budget-matched diffusion reinforcement learning on math and code. A Hugging Face daily-papers item dated 4 October 2026 instead reports a closed-loop activation-steering attack that uses a proportional-integral controller to steer denoising toward an adversary-chosen demographic answer. On ambiguous BBQ items, preference for the targeted group rose from 1.8 to 16.7 percentage points, with other demographic targets shifting by up to 37 points, while SocialStigmaQA stigmatizing answers rose from 17.6% to 58.1%. Constant-strength steering moved answers less and corrupted nearly three times as many outputs, and each attack took about 40 minutes on one GPU. The sources share no overlapping experimental claims on which they disagree.
- Pivot-SD (arXiv, 2026-10-02) is offline self-distillation for masked diffusion LMs that trains only on high-impact pivot commitments chosen by an uncertainty-reduction information-gain metric.
- Successful-trajectory pivots get cross-entropy; failed-trajectory pivots get unlikelihood loss; other tokens are left untouched.
- With 200 questions and four rollouts each, Pivot-SD improved LLaDA-8B-Instruct over full-sequence SFT and budget-matched diffusion RL on math and code benchmarks.
- A 2026-10-04 report describes a separate closed-loop attack: a proportional-integral controller adapts an activation steering vector toward an adversary-chosen demographic answer during denoising.
- On ambiguous BBQ questions, LLaDA-8B-Instruct targeted-group preference rose from 1.8 to 16.7 percentage points; other demographic targets shifted by up to 37 percentage points.
- On SocialStigmaQA, stigmatizing answers rose from 17.6% to 58.1%.
- Constant-strength steering shifted answers less and corrupted nearly three times as many outputs; each attack took about 40 minutes on one GPU.
Coverage timelineoldest first · each row is one article
- · 6d agoPivot-SD: Efficient Self-Distillation for Masked Diffusion Language Models
arXiv cs.AI / cs.LG / cs.CL· 40
Pivot-SD trains masked diffusion LLMs only on high-impact token commitments, improving LLaDA-8B-Instruct math and code results with just 200 questions.
- · 4d agoNoise Out, Bias In: Targeted Bias Injection in Diffusion Language Models via Closed-Loop Activation Steering
Hugging Face daily papers· 58
A closed-loop steering attack injects targeted demographic bias into diffusion language models during denoising.