PixelDense: Dense Prediction as Representation Alignment for Pixel Diffusion
PixelDense separates semantic and geometric alignment teachers, improving pixel diffusion quality and training speed.
PixelDense uses dense-prediction models as representation-alignment targets for pixel-space diffusion, splitting DINOv2 and SAM2 into a semantic stream and Depth Anything v2 and Metric3D v2 into a geometric stream, with a weight-space orthogonality penalty. A flat sum of all four teachers underperformed the best geometric teacher because semantic and geometric gradients competed. Teachers are frozen during training and dropped at inference. On PixelGen and DeCo, PixelDense improved GenEval, DPG-Bench, and HPS v2.1, lifting PixelGen-XXL GenEval Overall from 0.7927 to 0.8093, with up to 53.1% PQ gain and 36.0% depth AbsRel reduction in partial-noise probes.