arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37638cs.CV

针对对比视觉-语言模型的目标视觉反事实解释

Targeted Visual Counterfactual Explanations for Contrastive Vision-Language Model

Van Bach Nguyen, Jörg Schlötterer, Christin Seifer

首次发表
浏览论文内容

中文总结 AI 辅助

针对CLIP零样本分类,提出掩码引导的自适应反事实解释方法MACE,通过源或差异归因构建可编辑区域,结合潜在扩散修复与冻结CLIP指导,在四个数据集上实现高目标成功率与真实感,并揭示有效性与保留的权衡。

中文摘要 AI 辅助

当前针对对比视觉-语言模型(如CLIP)的解释方法主要识别重要区域,而未展示如何改变输入以获得目标预测。我们引入了掩码引导的自适应反事实解释(MACE),这是一种专门为CLIP零样本分类设计的目标视觉反事实方法。MACE从源归因或源-目标归因差异中构建可编辑区域,并仅在需要达到指定目标类别时扩展掩码。随后,潜在扩散修复模型修改所选区域,而冻结的CLIP模型提供修改指导并将剩余图像内容锚定到原始输入。我们在ImageNet、Food-101、Oxford Pets和CUB-200上评估了MACE。源掩码变体在所有四个数据集上实现了最高的目标top-1成功率,而差异掩码变体产生了最小的像素级和感知变化以及最佳的真实感分数。两种变体在相同生成骨干网络上均比仅使用Stable Diffusion的基线提高了接近度和真实感。这些结果表明,自适应掩码引导编辑能产生有效的CLIP反事实。它们进一步揭示了反事实有效性与源图像保留之间的权衡。

英文摘要

Current explanation methods for contrastive vision-language models such as CLIP mainly identify important regions without showing how to change the input in order to get a target prediction. We introduce Mask-guided Adaptive Counterfactual Explanations (MACE), a targeted visual counterfactual method designed specifically for CLIP zero-shot classification. MACE constructs an editable region from either source attribution or source-target attribution differences and expands the mask only when needed to reach a specified target class. A latent diffusion inpainting model then modifies the selected region, while a frozen CLIP model provides modification guidance and anchors the remaining image content to the original input. We evaluate MACE on ImageNet, Food-101, Oxford Pets, and CUB-200. The source-mask variant achieves the highest target top-1 success rate across all four datasets, while the difference-mask variant produces the smallest pixel level and perceptual changes and the best realism scores. Both variants improve proximity and realism over a Stable Diffusion-only baseline using the same generative backbone. These results show that adaptive mask-guided editing produces effective CLIP counterfactuals. They further reveal a tradeoff between counterfactual validity and source-image preservation.

发表机构

  • Marburg University(马尔堡大学)

机构由 AI 辅助整理,请以论文原文为准。

↑