Fréchet 距离度量的是什么?一种方向性分解
What Does Fréchet Distance Measure? A Directional Decomposition
- Agency for Defense Development(国防发展局)
- Harvard University(哈佛大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对 Fréchet 距离标量值缺乏可解释性的问题,提出方向性 Fréchet 距离,通过分解最优传输位移揭示差异来源,并在图像、视频和蛋白质任务中验证其解释能力。
AI中文摘要:
Fréchet 距离是跨领域评估生成模型的事实标准,在图像领域表现为 FID,在视频领域表现为 FVD。它将生成分布与参考分布之间的差异汇总为单个标量,数值越低通常被解释为生成质量越好。然而,这种标量视角可能掩盖比较背后的驱动因素。例如,在 COCO 数据集中,增加扩散采样步数会提升 ImageReward 分数,却使 FID 恶化(增大)。受此不一致现象的启发,我们旨在通过揭示差异所在位置来使 Fréchet 距离更具可解释性。为此,我们引入了方向性 Fréchet 距离,即最优传输位移在给定方向上的期望平方投影。在图像、视频和蛋白质案例研究中,我们发现少量可解释的方向解释了大部分距离。我们利用这些方向,以 CLIP 嵌入所代表的语义概念来解释 FID 的增大,量化 FVD 对逐帧外观的偏向,并重新审视蛋白质 FID 的解释。我们在该 https URL 开源了代码库。
英文摘要:
The Fréchet distance is a de facto standard for evaluating generative models across domains, appearing as FID for images and FVD for videos. It summarizes the discrepancy between generated and reference distributions in a single scalar, with lower values typically interpreted as better generation quality. However, this scalar view can obscure what drives the comparison. For example, in COCO dataset, increasing the number of diffusion sampling steps improves ImageReward scores yet worsens (increases) FID. Motivated by this mismatch, we seek to make the Fréchet distance more interpretable by uncovering where the discrepancy lies. To this end, we introduce directional Fréchet distance, the expected squared projection of the optimal transport displacement onto a given direction. Across our image, video, and protein case studies, we find that a small number of interpretable directions account for much of the distance. We use these directions to explain the FID increase in terms of semantic concepts represented by CLIP embeddings, quantify FVD's bias toward per-frame appearance, and revisit the interpretation of Protein FID. We open-source our codebase at https://github.com/yhlee-add/directional-fd.