JLD:通过雅可比透镜的感知距离
JLD: Perceptual Distance Through A Jacobian Lens
浏览论文内容
中文总结 AI 辅助
JLD利用冻结视觉编码器的雅可比矩阵构造固定度量张量,无需人类标签即可实现跨分辨率鲁棒的感知距离,在多个基准上超越现有方法并支持视频。
中文摘要 AI 辅助
图像压缩、恢复和生成都需要一种方法来衡量两幅图像在人眼中的差异程度。像素误差忽略了人的视觉感知方式,而最准确的感知距离通常拟合人类判断,从而受限于固定的数据和分辨率。例如,当图像分辨率加倍时,DISTS与TID2013上人类评分的相关性从0.815降至0.717。我们引入了雅可比透镜距离(JLD),其感知几何源自冻结的视觉编码器,而非人类标签。JLD结合了早期补丁特征的局部性与后期编码器表示所捕获的感知敏感性。具体而言,我们使用编码器雅可比矩阵来识别早期特征空间中能最强影响编码器输出的方向,从而产生一个固定的度量张量E[J^T J],我们称之为雅可比透镜。该透镜仅从100张无标签图像拟合一次,耗时约35秒。在局部,这种构造在像素空间中定义了拉回度量,赋予JLD清晰的几何解释,可直接在真实图像上进行分析。在四个标准感知数据库中,JLD达到了最先进的性能,并持续优于LPIPS、DISTS、PieAPP和DreamSim。JLD对图像分辨率变化也具有鲁棒性,在TID2013上,当分辨率加倍时,其透镜项相关性几乎保持不变,仅从0.850降至0.845。我们进一步引入了JLD-fast,它比LPIPS-VGG快4倍,同时达到0.911的平均相关性。最后,JLD自然扩展到视频,在Waterloo IVC 4K上达到0.786的相关性,而VMAF为0.611。
英文摘要
Image compression, restoration, and generation all require a way to measure how different two images look to a person. Pixel error ignores how people see, while the most accurate perceptual distances are typically fitted to human judgments, tying them to a fixed data and resolution. For example, when image resolution is doubled, the correlation of DISTS with human scores on TID2013 drops from 0.815 to 0.717. We introduce the Jacobian Lens Distance (JLD), which derives its perceptual geometry from a frozen vision encoder rather than from human labels. JLD combines the locality of early patch features with the perceptual sensitivity captured by later encoder representations. Specifically, we use the encoder Jacobian to identify directions in the early feature space that most strongly affect the encoder output, producing a fixed metric tensor, $E[J^\top J]$, which we call the Jacobian lens. The lens is fitted only once from 100 unlabeled images, taking about 35 seconds. Locally, this construction defines a pullback metric in pixel space, giving JLD a clear geometric interpretation that can be directly analyzed on real images. Across four standard perceptual databases, JLD achieves state-of-the-art performance and consistently outperforms LPIPS, DISTS, PieAPP, and DreamSim. JLD is also robust to changes in image resolution, on TID2013, its lens-term correlation remains nearly unchanged when the resolution is doubled, decreasing only from 0.850 to 0.845. We further introduce JLD-fast, which is $4\times$ faster than LPIPS-VGG while achieving a mean correlation of 0.911. Finally, JLD naturally extends to video, reaching a correlation of 0.786 on Waterloo IVC 4K compared with 0.611 for VMAF.
发表机构
- The University of Texas at Austin(德克萨斯大学奥斯汀分校)
- Google(谷歌)
- University of Colorado Boulder(科罗拉多大学博尔德分校)
机构由 AI 辅助整理,请以论文原文为准。