发表机构
Jagiellonian University; University of Information Technology and Management in Rzeszów; Jagiellonian University Medical College(雅盖隆大学; 热舒夫信息技术与管理大学; 雅盖隆大学医学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究提出基于充分性的显著性方法SEAMS,通过简单优化框架,利用软掩码等实现对冻结模型输出的保留,能生成多种解释,掩码紧凑可解释,且不同架构依赖证据不同,突出视觉解释的架构依赖性。
AI 中文摘要
显著性图在识别足以保留模型行为的图像区域时最为有用。我们引入了SEAMS,一种基于充分性的显著性方法,它使用保留目标直接优化软掩码。给定一个冻结的可微模型输出,如类别概率、CLS嵌入或令牌表示,SEAMS搜索一个紧凑的掩码来保留所选输出。该方法依赖于一个基于软掩码、可学习预算和完全由查询图像生成的三向图像合成的简单优化框架。实验表明,相同的优化管道通过仅更改保留目标,就能生成对象级、类别条件和令牌级解释。生成的掩码紧凑、可解释、在随机初始化下稳定,且在插入和删除基准测试中具有竞争力。我们的结果还表明,不同架构在实现相似保留保真度时通常依赖不同的充分证据,突出了视觉解释的架构依赖性。
英文摘要
Saliency maps are most useful when they identify the image regions that are sufficient to preserve a model's behaviour. We introduce SEAMS, a sufficiency-based saliency method that directly optimises a soft mask using a preservation objective. Given a frozen differentiable model output, such as a class probability, CLS embedding, or token representation, SEAMS searches for a compact mask that preserves the selected output. The approach relies on a simple optimisation framework based on soft masks, a learnable budget, and a three-way image composite generated entirely from the query image. As a result, it requires no auxiliary distractor dataset, architecture-specific attribution mechanism, or differentiable top-k relaxation. Experiments with frozen ViT-S/16 and ConvNeXt models show that the same optimisation pipeline can generate object-level, class-conditioned, and token-level explanations by changing only the preserved target. The resulting masks are compact, interpretable, stable across random initialisations, and competitive on insertion and deletion benchmarks. Our results also indicate that different architectures often rely on different sufficient evidence while achieving similar preservation fidelity, highlighting the architecture-dependent nature of visual explanations.