发表机构
School of Cyber Science and Technology, Shenzhen Campus of Sun Yat-sen University; University of Chinese Academy of Sciences; School of Automation, Nanjing University of Information Science and Technology; Department of Mechanical Engineering, Imperial College London; College of Computing and Data Science, Nanyang Technological University(中山大学深圳校区网络科学与技术学院; 中国科学院大学; 南京信息工程大学自动化学院; 伦敦帝国理工学院机械工程系; 南洋理工大学计算与数据科学学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对归因方法在几何变换下的不足,提出基于图像区域子模搜索的无注释归因正则化框架,通过子模排序损失正则化归因及证据选择过程,实验表明该方法能提升归因稳定性且性能损失小。
AI 中文摘要
归因方法被广泛用于刻画模型预测背后的证据,但在改善模型行为方面的潜力尚未充分探索。标签保持几何变换下的归因不一致可能表明对变换敏感的证据依赖,这促使进行归因正则化。然而,这种监督仅在归因忠实地反映驱动预测的证据时才有效。现有自监督方法通常对齐基于梯度的映射,其有限的忠实性意味着归因一致性并不一定意味着潜在决策过程的一致性,变换稳健性问题仍未解决。我们提出了一个基于图像区域子模搜索的无注释归因正则化框架。通过测量候选子集如何影响模型输出,搜索提取紧凑、类区分性的证据作为搜索衍生的监督。我们还引入了具有路径一致性和终止对齐项的子模排序损失,分别沿着配对搜索轨迹对齐空间对应候选排序,并鼓励变换轨迹在目标终端步骤满足停止标准。该损失为正则化最终归因和离散的证据选择过程提供了可微替代。在ImageNet-100上的实验表明,我们的方法显著提高了归因稳定性,在ViT-B/16上的插入和删除操作仅导致0.28点的准确率下降,在ViT-L/16上也有类似收益。在ImageNet-1K上,它提高了ResNet-50和ConvNeXt-B上变换输入的准确率,同时将干净准确率下降限制在0.30点,证明了在性能损失最小的情况下更一致的证据依赖。代码即将发布。
英文摘要
Attribution methods are widely used to characterize the evidence underlying model predictions, yet their potential to improve model behavior remains underexplored. Attribution inconsistency under label-preserving geometric transformations may indicate transformation-sensitive evidence reliance, motivating attribution regularization. However, such supervision is valid only when attribution faithfully reflects the evidence driving predictions. Existing self-supervised methods typically align gradient-based maps such as Grad-CAM, whose limited faithfulness means that attribution consistency need not imply consistency of the underlying decision process, leaving transformation robustness unresolved. We propose an annotation-free attribution regularization framework based on submodular search over image regions. By measuring how candidate subsets affect model outputs, the search extracts compact, class-discriminative evidence as search-derived supervision. We further introduce a submodular ranking loss with path-consistency and termination-alignment terms that respectively align spatially corresponding candidate rankings along paired search trajectories and encourage the transformed trajectory to satisfy the stopping criterion at the target terminal step. The loss provides a differentiable surrogate for regularizing both final attributions and the otherwise discrete evidence-selection process. Experiments on ImageNet-100 show that our method substantially improves attribution stability, Insertion, and Deletion on ViT-B/16 with only a 0.28-point accuracy drop, with similar gains on ViT-L/16. On ImageNet-1K, it improves transformed-input accuracy on ResNet-50 and ConvNeXt-B while limiting the clean-accuracy drop to 0.30 points, demonstrating more consistent evidence reliance with minimal performance loss. Code will be released soon.