发表机构
Hertie Institute for AI in Brain Health, University of Tübingen; Charité - Universitätsmedizin Berlin; Humboldt-Universität zu Berlin; Tübingen AI Center, University of Tübingen; German Center for Mental Health (DZPG); Fraunhofer Heinrich Hertz Institute; Technische Universität Berlin(蒂宾根大学赫蒂脑健康人工智能研究所; 柏林夏里特医学院; 柏林洪堡大学; 蒂宾根大学蒂宾根人工智能中心; 德国心理健康中心; 弗劳恩霍夫海因里希赫兹研究所; 柏林工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对深度神经网络的捷径学习问题,提出ICON分解方法,通过量化概念在考虑其他概念及结果后的方差解释度,在合成数据及皮肤病变、脑成像模型上实现更准确的概念级解释,可用于模型审计。
AI 中文摘要
深度神经网络常利用训练数据中的虚假关联,这种缺陷被称为捷径学习。基于概念的可解释性方法通过测试患者性别或扫描仪设置等概念能否从网络层解码来筛查捷径。由于每个概念是单独评估的,这些方法可能误将概念间的相关性当作模型使用它们的证据。我们提出ICON分解,它会在考虑所有其他概念和结果后,量化每个概念对某一层方差的解释程度。在具有已知真实值的合成数据上,ICON比7种替代基线方法更准确地恢复概念重要性;在皮肤病变和脑成像模型上,它能分离模型真正依赖的概念、量化所有提供概念无法解释的表征,并生成稀疏解释,我们通过重训练和分布外测试验证了这些解释。
英文摘要
Deep neural networks often exploit spurious associations, a failure known as shortcut learning. Before deployment, models should be audited for reliance on a set of concepts, such as acquisition artifacts or demographics. Current methods, such as linear probes and concept activation vectors, measure reliance by asking whether each concept, in isolation, is decodable from a layer. Their scores therefore reflect not only reliance but also correlations in the audit dataset. We introduce Independent Canonical cONcept (ICON) decomposition, which quantifies the share of a layer's variance each concept explains, conditional on all other concepts and the outcome. ICON scores are variance shares, comparable across layers and between continuous and categorical concepts. ICON also reports the share the set leaves unexplained. On simulated data, ICON recovers the true importance more accurately than seven baselines. On skin-cancer and neuroimaging models, ICON distinguishes learned shortcuts from correlated concepts, confirmed by retraining and out-of-distribution tests.
Comments44 pages, 12 figures, 3 tables. Includes Extended Data (7 figures, 2 tables). Code: https://github.com/RoshanRane/ICON_decomposition