发表机构
Bielefeld University(比勒费尔德大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出对比式概念重要性(CCI)方法,将目标与对比类别间的logit边际归因于自动提取的视觉概念,可区分概念对类别对的不同效应,经ImageNet实验证实其能揭示普通概念重要性未捕捉的特定类别对模型行为。
AI 中文摘要
基于概念的解释是解释复杂黑箱模型决策的流行方式,通过语义上有意义、人类可理解的概念实现。为将这些概念的贡献归因于模型决策,特征归因方法用于量化每个概念对模型输出的贡献程度。这些归因通常针对单个输出类别计算,因此回答的是非对比性的“为什么是P?”的问题。但在许多情况下,例如误分类、类别混淆和低边际预测时,更自然的问题是“为什么是P而不是Q?”。我们引入对比式概念重要性(CCI),将目标类别与对比(或负)类别之间的logit边际归因于自动提取的视觉概念基中的概念。得到的分数是带符号的,表明概念是支持目标类别而非对比类别,还是支持对比类别而非目标类别,并且可以分解为目标logit和对比logit效应。这使得能够区分全局重要概念与特定影响类别对区分的概念,包括它们的效应是否共享、单侧或直接对比。我们在ImageNet类别对上使用CRAFT风格的概念基、插入和删除曲线、逐logit分解分析以及语义类别层次结构评估该方法。结果表明,对比式概念重要性揭示了普通概念重要性无法捕捉的特定类别对的模型行为,并且可以根据语义超类结构评估高度对比的概念,以确定它们是否影响细粒度区分而非宽泛类别证据。
英文摘要
Concept-based explanations are a prevalent way to explain the decisions of complex black-box methods through semantically meaningful, human interpretable concepts. To attribute the contribution of such concepts to a model's decisions, feature attribution methods are used to quantify how strongly each concept contributes to a model output. These attributions are typically computed for a single output class and therefore answer a non-contrastive "why P?" question. In many situations, however, such as cases of misclassification, class confusion, and low- margin predictions, the more natural question to ask is "why P rather than Q?". We introduce contrastive concept importance, which attributes the logit margin between a target class and a contrast, or foil, class to concepts in an automatically extracted visual concept basis. The resulting scores are signed, indicating whether a concept supports the target over the foil or the foil over the target, and can be decomposed into target-logit and foil-logit effects. This makes it possible to distinguish globally important concepts from concepts that specifically influence a class-pair distinction, including whether their effect is shared, one-sided, or directly contrastive. We evaluate our method both qualitatively and quantitatively on a range of ImageNet class pairs. Our results show that contrastive concept importance reveals class-pair specific model behavior that is not captured by standard concept importance alone, as well as capturing information on the semantic structure of the underlying ImageNet classes.