基于扩散模型的基于概念的视觉反事实解释
Concept-based Visual Counterfactual Explanations with Diffusion Models
浏览论文内容
中文总结 AI 辅助
研究视觉反事实解释问题,提出C-VCE框架,通过概念瓶颈层将分类器融入生成模型,由概念引导反事实解释,在采样时可切换概念,添加正则化器和掩码控制编辑,在基准测试中表现良好,是实用工具且为模型理解和使用提供新思路。
中文摘要 AI 辅助
视觉反事实解释旨在回答‘对该图像进行何种最小更改会改变模型预测?’,在安全关键领域愈发重要。现有基于扩散的方法依赖外部分类器,在处理噪声图像时不可靠且难以部署。我们引入C-VCE,通过概念瓶颈层将分类器直接构建到生成模型中,使反事实解释由人类可解释的特征(概念)引导。在采样期间用户可切换语义概念,最小化调整相关图像区域并保留其余部分。添加正则化器和梯度掩码控制编辑,在CelebA等基准测试中,C-VCE匹配或提高翻转率,生成的反事实图像视觉上更接近输入且失真更小。这使其成为视觉系统实用工具,更广泛地说,结果表明暴露和控制内部概念层是使强大生成模型更易理解和安全使用的有效途径。
英文摘要
Visual counterfactual explanations aim to answer "what minimal change to this image would flip the model's prediction?", and are increasingly important as vision models are deployed in safety-critical domains (e.g., medicine). Existing diffusion-based methods can produce realistic edits, but they rely on external classifiers that must work reliably on noisy images, which makes them fragile and hard to deploy for robust explanations. We introduce C-VCE, a new diffusion framework that builds the classifier directly into the generative model via a concept bottleneck layer, so that counterfactuals are guided by human-interpretable features (concepts) instead of a separate noise robust classifier that works with pixel-level edits. Our model lets users to toggle on/off semantic concepts during sampling, then minimally adjusts relevant image regions, while preserving the rest of the image, respecting feature correlations. To keep edits small and controlled, we add a simple probabilistic regularizer that balances "change the prediction" against "stay close to the original", plus a gradient-based mask that confines modifications to the most relevant regions. On benchmarks such as CelebA, C-VCE matches or improves flip rates while producing counterfactuals that are visually closer to the input and less distorted than baselines that depend on separate noisy-image classifiers. These properties make C-VCE a practical tool for vision systems where users need concrete "what-if" images without having to trust an additional, noise-robust classifier. More broadly, our results suggest that exposing and controlling an internal concept layer is a promising way to make powerful generative models easier to understand and safer to use.
发表机构
- Università della Svizzera italiana(瑞士意大利语区大学)
机构由 AI 辅助整理,请以论文原文为准。