Aligning One-Step Generative Models with Reward-Weighted Transport Distillation
RWTD post-trains one-step generators, raising SANA Sprint GenEval from 0.73 to 0.80.
Reward-Weighted Transport Distillation (RWTD) post-trains one-step generative models using only generated samples and scalar reward evaluations. It builds an adaptive target that mixes separately reward-tilted current and reference distributions, implemented with feature-space optimal transport and fixed-point regression. The analysis says its fixed points interpolate between off-policy tilting of the reference and on-policy tilting of the current model. Empirically, RWTD raised the GenEval score of the one-step SANA Sprint 1.6B backbone from 0.73 to 0.80, with preference experiments showing cross-reward generalization and preserved compositional ability.
- Post-training uses only generated samples and scalar rewards
- Adaptive target mixes tilted current and reference distributions
- SANA Sprint 1.6B GenEval rises from 0.73 to 0.80
- Preference tests show cross-reward gains with composition preserved
Full article175 words · extracted from huggingface.co · click to collapse
One-step generators enable high-quality visual generation with a single network evaluation, but their post-training is difficult: general implicit generators provide neither tractable likelihoods nor denoising trajectories, and many rewards are non-differentiable. We introduce Reward-Weighted Transport Distillation (RWTD), a post-training method that requires only generated samples and scalar reward evaluations. Rather than aligning solely to the conventional reward-tilted reference distribution, RWTD constructs an adaptive target that mixes separately tilted current and reference distributions. The current component incorporates improvements discovered during training, while the reference component anchors the target to the pretrained generator. RWTD realizes this target through feature-space optimal transport and fixed-point regression. Theoretical analysis shows that the fixed-point distributions of RWTD interpolate between off-policy reward tilting of the reference and on-policy tilting of the current model, providing a principled approach to balancing reward adaptation with retention of prior knowledge. Empirically, RWTD substantially improves the GenEval score of the one-step SANA Sprint 1.6B backbone from 0.73 to 0.80, while separate preference alignment experiments demonstrate strong cross-reward generalization that yields balanced improvements and preservation of compositional capabilities.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.30840