IDRF: Inverse-Distilled Reward Fine-tuning of Masked Discrete Diffusion Models
IDRF reward-fine-tunes few-step masked discrete diffusion models via inverse-distillation regularization, achieving high reward with up to 32x fewer denoising steps while mitigating reward hacking.
IDRF is a reward fine-tuning framework for few-step masked discrete diffusion generators that replaces the intractable sequence-level KL penalty with an inverse-distillation regularization, proven to upper-bound the sequence-level KL divergence to the reference distribution given an optimal auxiliary denoiser. It optimizes a trajectory-based surrogate without reference-model rollouts, treating few-step generation as a finite-horizon MDP with a clipped policy-gradient objective. Across DNA, image, and text generation, IDRF achieves high reward with up to 32x fewer denoising steps than the reference while mitigating reward hacking and preserving sample quality.
- Inverse-distillation regularization provably upper-bounds sequence-level KL to reference distribution
- Few-step generation framed as finite-horizon MDP with clipped policy-gradient optimization
- No reference-model rollouts needed; student keeps its own few-step sampler
- High reward with up to 32x fewer denoising steps on DNA, image, and text tasks
- Mitigates reward hacking while preserving sample quality
Full article144 words · extracted from arxiv.org · click to collapse
Masked discrete diffusion models offer a promising alternative to autoregressive generation, but iterative sampling can be costly, and intractable sequence likelihoods complicate reward fine-tuning. We introduce IDRF, a framework for reward fine-tuning of few-step masked discrete diffusion generators. Starting from a standard reverse-KL-regularized objective, IDRF replaces the intractable sequence-level KL penalty with inverse-distillation regularization. With an optimal auxiliary denoiser, we prove that the population inverse-distillation loss upper-bounds the sequence-level KL divergence to the reference distribution. IDRF optimizes a trajectory-based surrogate of this loss without reference-model rollouts, so the student keeps its own few-step sampler. We view few-step generation as a finite-horizon Markov decision process and optimize reward with a clipped policy-gradient objective over the student's trajectories. Across DNA, image, and text generation, IDRF achieves high reward with up to $32\times$ fewer denoising steps than the reference while mitigating reward hacking and preserving sample quality.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2610.03641