ReGain: Restoring Subject Fidelity in Personalization on Synthetic Images
ReGain downscales inflated guidance so DreamBooth on synthetic images recovers much of real-photo subject fidelity.
DreamBooth personalization on diffusion-generated subject images degrades fidelity, producing oversaturated color and excess high-frequency detail. The authors trace this to classifier-free guidance: the angle and difference norm between conditional and unconditional noise predictions inflate, especially at high frequencies and nearby prompts. ReGain is a training-free sampling correction that downscales frequency bands whose guidance is inflated relative to the base model. On Stable Diffusion v1.5 it closes 51-64% of the gap to a real-photo personalization on DINO, DINOv2, and CLIP-I, and it also helps SDXL and SD 3.5 without hurting text alignment.
- DreamBooth on synthetic subject images oversaturates color and adds excess high-frequency detail.
- Cause is inflated classifier-free guidance versus the same recipe on real photos.
- ReGain scales inflated frequency bands at sampling time and needs no real photos.
- On SD 1.5 it closes 51-64% of the DINO, DINOv2, and CLIP-I fidelity gap.
- Gains also appear on SDXL and SD 3.5 while text alignment is preserved.
Full article235 words · extracted from huggingface.co · click to collapse
Text-to-image diffusion models are personalized to a subject by DreamBooth fine-tuning on a handful of its images. Increasingly, these images come from a diffusion model rather than a camera. We show that fine-tuning on such synthetic images degrades subject fidelity, producing oversaturated color and excess high-frequency detail. To isolate the cause, we fine-tune two models from the same base model with the same DreamBooth recipe, one on real photos of a subject and one on synthetic images of that subject generated by the first. We trace the degradation to classifier-free guidance (CFG). For the model personalized on synthetic images, the angle between the conditional and unconditional noise predictions, and with it the norm of their difference, is much larger than for the model personalized on real photos. This inflation grows toward high frequencies and also appears at other prompts semantically close to the subject, such as its class noun, but not at unrelated ones. We propose ReGain, a training-free correction applied at sampling time that measures how much each frequency band of the guidance is inflated relative to the base model and scales that band down accordingly. ReGain needs no real photos. On Stable Diffusion v1.5, ReGain closes 51-64% of the subject-fidelity gap to the model personalized on real photos, as measured by DINO, DINOv2 and CLIP-I. It also improves subject fidelity on SDXL and SD 3.5 and preserves text alignment on all three backbones.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.38680