JLD: Perceptual Distance Through A Jacobian Lens
Jacobian Lens Distance measures perceptual image difference from a frozen vision encoder without human-label fitting.
Jacobian Lens Distance (JLD) derives a perceptual image metric from a frozen vision encoder rather than human judgments, fitting a Jacobian-based metric tensor once on 100 unlabeled images in about 35 seconds. Across four perceptual databases it outperforms LPIPS, DISTS, PieAPP, and DreamSim, and its TID2013 correlation barely changes when resolution doubles, from 0.850 to 0.845, unlike DISTS which falls from 0.815 to 0.717. JLD-fast is four times faster than LPIPS-VGG at 0.911 mean correlation, and the metric reaches 0.786 correlation on Waterloo IVC 4K video versus 0.611 for VMAF.
- JLD builds a fixed metric tensor from a frozen encoder Jacobian using 100 unlabeled images.
- On TID2013, correlation stays nearly flat when resolution doubles, 0.850 to 0.845.
- JLD-fast is 4× faster than LPIPS-VGG with mean correlation 0.911.
- On Waterloo IVC 4K video, JLD scores 0.786 versus 0.611 for VMAF.
Full article259 words · extracted from huggingface.co · click to collapse
Image compression, restoration, and generation all require a way to measure how different two images look to a person. Pixel error ignores how people see, while the most accurate perceptual distances are typically fitted to human judgments, tying them to a fixed data and resolution. For example, when image resolution is doubled, the correlation of DISTS with human scores on TID2013 drops from 0.815 to 0.717. We introduce the Jacobian Lens Distance (JLD), which derives its perceptual geometry from a frozen vision encoder rather than from human labels. JLD combines the locality of early patch features with the perceptual sensitivity captured by later encoder representations. Specifically, we use the encoder Jacobian to identify directions in the early feature space that most strongly affect the encoder output, producing a fixed metric tensor, E[J^top J], which we call the Jacobian lens. The lens is fitted only once from 100 unlabeled images, taking about 35 seconds. Locally, this construction defines a pullback metric in pixel space, giving JLD a clear geometric interpretation that can be directly analyzed on real images. Across four standard perceptual databases, JLD achieves state-of-the-art performance and consistently outperforms LPIPS, DISTS, PieAPP, and DreamSim. JLD is also robust to changes in image resolution, on TID2013, its lens-term correlation remains nearly unchanged when the resolution is doubled, decreasing only from 0.850 to 0.845. We further introduce JLD-fast, which is 4times faster than LPIPS-VGG while achieving a mean correlation of 0.911. Finally, JLD naturally extends to video, reaching a correlation of 0.786 on Waterloo IVC 4K compared with 0.611 for VMAF.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2610.05967