arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.14051cs.CV

用于审计视觉模型偏差的多个捷径组的发现与空间特征表征

Discovery and Spatial Characterisation of Multiple Shortcut Groups for Auditing Vision Model Bias

Akshit Achara, Vishnunarayan Manickam, Thomas Day, Esther Puyol Anton, Alexander Hammers, Andrew P. King

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对视觉模型偏差审计,通过K-means和非负矩阵分解发现图像中重复的捷径空间模式,结合特征干预方法降低模型性能差异,适用于多个数据集与模型。

中文摘要 AI 辅助

在存在虚假关联的数据集上训练的深度学习模型可实现较高的平均准确率,但其依赖的捷径特征无法泛化到分布外场景。虽然分布外测试能凸显由捷径学习导致的子群体性能差异,但无法定位图像中与捷径学习相关的区域。现有研究大多使用可解释性方法的归因图来理解虚假关联的空间特性,例如条件对齐方法通过对比任务模型、敏感属性模型及去偏参考模型的归因图,将任务相关证据与虚假关联绑定的证据分离,为每张图像生成与捷径对齐及与任务对齐的贡献图。然而,现有方法会在整个数据集上聚合这些图,可能掩盖仅在部分图像子集中出现的重复空间捷径模式。我们通过K-means和非负矩阵分解将单张图像的捷径与任务贡献图分组为重复空间模式,并通过贡献图和代表性示例可视化所得的捷径组。在CelebA、CheXpert、Waterbirds、Camelyon17和ISIC2019数据集,以及ResNet和ViT模型上,所发现的捷径组揭示了捷径与任务贡献的共享及独特空间模式,具有不同的子群体组成和错误率,支持对错误率较高的图像子集进行针对性检查。我们通过输入遮挡和内部测试时干预实验表明,屏蔽或抑制任务贡献区域会大幅降低模型分类性能,并提出一种结合捷径抑制与任务放大的特征干预方法,该方法可普遍降低性能差异。

英文摘要

Deep learning models trained on datasets with spurious correlations can achieve high average accuracy whilst relying on shortcut features that do not generalise out of distribution. Whilst out-of-distribution testing highlights subgroup performance disparities arising from shortcut learning, it does not localise the regions within images that are associated with it. Existing research mostly uses attribution maps from interpretability methods to understand the spatial nature of spurious correlations. For example, conditional alignment methods separate task-relevant evidence from evidence tied to spurious correlations by comparing attribution maps from a task model, a sensitive attribute model, and a bias-reduced reference model. This yields shortcut-aligned and task-aligned contribution maps for each image. However, existing methods aggregate these maps across the dataset, potentially masking recurring spatial shortcut patterns that occur only in subsets of images. We address this limitation by grouping per-image shortcut and task contribution maps into recurring spatial patterns using K-means and non-negative matrix factorisation, and visualising the resulting shortcut groups through contribution maps and representative examples. Across CelebA, CheXpert, Waterbirds, Camelyon17, and ISIC2019, and across ResNet and ViT models, the discovered shortcut groups reveal both shared and distinct spatial patterns of shortcut and task contribution, with varying subgroup composition and error rates, enabling targeted inspection of image subsets with higher error rates. We perform input occlusion and internal test-time interventions to show that masking or suppressing task contribution regions substantially degrades the model classification performance and propose a combined shortcut suppression and task amplification feature intervention approach which generally reduces performance disparities.

发表机构

  • King’s College London(伦敦国王学院)
  • School of Biomedical Engineering and Imaging Sciences(生物医学工程与影像科学学院)

机构由 AI 辅助整理,请以论文原文为准。

↑