AI 中文总结
该研究提出与模型和度量无关的框架,用稀疏自编码器提取的概念激活向量解释图像相似度,通过潜在扰动验证其符合数据分布,可用于理解模型的相似度判断。
AI 中文摘要
图像相似度是诸多计算机视觉应用的基础,但通常难以明确两张图像为何会获得高或低的相似度分数。现有可解释性方法常依赖基于梯度的归因图为相似度提供局部解释,这些方法难以对嵌入空间中驱动相似度的具体因素(如纹理、形状或颜色)提供全局见解。我们提出一种与模型和度量无关的框架,该框架使用通过稀疏自编码器(SAEs)自动提取的概念激活向量(CAVs)来解释图像相似度。给定一对图像,我们沿已发现的概念方向对其嵌入进行扰动,并测量所选相似度函数的变化,从而得到概念重要性。对于图像对,我们提供带有概念归因图的定位。我们将此过程扩展到组级设置,解释是什么驱动了一组图像而非单个图像对之间的相似度,此外,我们引入了示例检索,旨在恢复具有相似相似度贡献原因的样本。我们的实验表明,我们的潜在扰动比像素空间基线更符合潜在数据分布,且概念重要性可线性恢复真实相似度分数。定性结果进一步证实了我们的方法在理解模型的个体和组相似度判断方面的实用性。
英文摘要
Image similarity underlies many computer vision applications, yet it is often unclear why two images receive a high or low similarity score. Existing explainability methods often rely on gradient-based attribution maps to provide local justifications for similarity. These approaches struggle to provide global insights into what specifically drives similarity in regions of an embedding space, such as texture, shape, or color. We introduce a model- and metric-agnostic framework that explains image similarity using Concept Activation Vectors (CAVs) extracted automatically via Sparse Autoencoders (SAEs). Given a pair of images, we perturb their embeddings along discovered concept directions and measure the resulting change in a chosen similarity function, yielding concept importances. For image pairs, we provide localization with concept attribution maps. We extend this procedure to group-level settings, explaining what drives similarity across a cluster of images rather than a single pair, and further, we introduce Exemplar Retrieval, aiming to recover samples with similar reasons contributing to similarity. Our experiments show that our latent perturbations are more faithful to the underlying data distribution than pixel-space baselines, and that concept importances linearly recover the true similarity score. Qualitative results further confirm the usefulness of our methods in understanding a model's individual and group similarity judgments.