FID隐藏了什么:检测、排序和诊断生成式评估中的偏差
What FID Hides: Detecting, Ranking, and Diagnosing Deviations in Generative Evaluation
查看机构详情
- University of California, Davis(加利福尼亚大学戴维斯分校)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
该研究针对生成式评估指标FID的缺陷,提出ZID方法,可检测、排序并诊断生成式模型的分布偏差,还能区分欠/过离散,在FID失效场景下有效。
中文摘要 AI 辅助
生成式模型通常通过Fréchet Inception距离(FID)和Kernel Inception距离(KID)进行排序,但FID的前两阶矩摘要可能遗漏分布差异,且仅报告标量间隙本身并非针对采样变异的校准检验。FID的矩限制有具体后果:在ImageNet上,仅优化以匹配参考Inception均值和协方差的视觉不可识别图像,获得的FID为24.7,而保留的真实图像FID为58.6(数值越低越好)。此外,FID和KID是标量差异,在两个样本交换时保持不变,因此不编码离散度变化的方向:是模式崩溃时可能出现的欠离散,还是过离散。我们引入ZID(Z-resolved Integrated Diagnostic),它结合了来自秩图(RISE)和高斯核(两种带宽下的GPK)的六个标准化位置和离散度敏感量。ZID不要求单个标量承担不兼容的角色,而是报告三个关联输出:用于排序偏差幅度的指数、用于检验分布相等性的排列p值,以及用于诊断的带符号离散度读数。在受控实验中,ZID可检测广泛的偏差,其分数沿相应扫描跟踪严重程度的增加,包括FID保持平稳或反转的情况。在DiT-XL/2和SiT-XL/2引导扫描中,ZID检测到与真实数据的偏差,其带符号读数将高引导多样性崩溃标记为欠离散。
英文摘要
Generative models are commonly ranked by Fréchet Inception Distance (FID) and Kernel Inception Distance (KID), yet FID's first-two-moment summary can miss distributional differences, and a reported scalar gap alone is not a calibrated test against sampling variation. FID's moment restriction has concrete consequences: on ImageNet, visually unrecognizable images optimized only to match the reference Inception mean and covariance obtain FID $24.7$ versus $58.6$ for held-out real images (lower is better). Moreover, FID and KID are scalar discrepancies that are unchanged when the two samples are exchanged and therefore do not encode the direction of a dispersion change: under-dispersion, as can occur in mode collapse, versus over-dispersion. We introduce \textbf{ZID} (\emph{Z-resolved Integrated Diagnostic}), which combines six standardized location- and dispersion-sensitive arms from a rank graph (RISE) and Gaussian kernels (GPK at two bandwidths). Rather than asking one scalar to serve incompatible roles, ZID reports three linked outputs: an index for ranking departure magnitude, a permutation $p$-value for testing distributional equality, and a signed dispersion readout for diagnosis. In controlled experiments, ZID detects a broad range of departures, and its score tracks increasing severity along the corresponding sweeps, including cases in which FID is flat or reversed. On DiT-XL/2 and SiT-XL/2 guidance sweeps, ZID detects departure from real data, and its signed readout labels the high-guidance diversity collapse as under-dispersion.