arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.10301cs.LGstat.ML

重新审视可解释人工智能:通过模型无关的概念字典

Revisiting Explainable AI through Model-Independent Concept Dictionaries

Thomas Schnake, Doreen Schöppenthau, Alexander Meyer, Jacques Corbeil, Klaus-Robert Müller, Grégoire Montavon

首次发表
浏览论文内容

中文总结 AI 辅助

针对现有XAI方法依赖可解释特征或架构特定内部抽象的问题,提出DictXAI,通过输入域中的预定义概念字典和稀疏编码实现模型无关的可解释归因,可识别AI故障并提升人机对齐,优于经典XAI。

中文摘要 AI 辅助

现代人工智能的应用依赖于日益复杂的模型。可解释人工智能(XAI)已成为一组旨在提高模型透明度的技术。然而,现有的XAI方法通常假设输入特征本身是可解释的,或者依赖于难以表征且高度特定于架构的中间内部抽象,这阻碍了跨模型的一致使用。为了解决这些局限性,我们提出了DictXAI,一种通过字典直接在输入域中定义概念的方法——一个大型的、可能过完备的预定义元素集合,每个元素都具有可解释的含义。在技术上,DictXAI首先计算输入的稀疏编码,然后将模型的预测归因于相关的字典元素。我们展示了DictXAI解释的可操作性,表明它们可以将人工智能故障(例如,聪明的汉斯效应)直接归因于数据中可识别的伪影模式,同时促进人类与人工智能在复杂生物医学信号上的对齐。我们进一步展示了我们的方法在各种字典上的操作能力,包括学习到的图像基、用于心电图的解析定义波形以及实验获取的字典元素。总体而言,我们的结果表明,与经典的XAI或现有的基于概念的方法相比,DictXAI提供了更可解释、可操作且架构无关的见解。

英文摘要

Modern applications of AI rely on increasingly complex models. Explainable AI (XAI) has emerged as a set of techniques aimed at improving model transparency. However, existing XAI methods typically assume input features to be inherently interpretable, or they rely on intermediate internal abstractions that are difficult to characterize and highly architecture-specific, hindering consistent use across models. To address these limitations, we propose DictXAI, a method that defines concepts directly in the input domain via a dictionary---a large, potentially overcomplete set of predefined elements, each carrying an interpretable meaning. Technically, DictXAI first computes a sparse code of the input and then attributes the model's prediction to the associated dictionary elements. We demonstrate the actionable nature of DictXAI explanations, showing that they can attribute AI malfunctions (e.g., Clever Hans effects) directly to identifiable artifact patterns in the data, while fostering human-AI alignment on intricate biomedical signals. We further demonstrate our method's ability to operate across a wide variety of dictionaries, including learned image bases, analytically defined waveforms for electrocardiography, and experimentally acquired dictionary elements. Overall, our results show that DictXAI provides more interpretable, actionable, and architecture-agnostic insights than classical XAI or existing concept-based approaches.

发表机构

  • University of Toronto(多伦多大学)
  • Vector Institute for Artificial Intelligence(向量人工智能研究所)
  • BIFOLD – Berlin Institute for the Foundations of Learning and Data(BIFOLD – 柏林学习与数据基础研究所)
  • Institute for AI in Medicine (IKIM), Charité – Universitätsmedizin Berlin(柏林夏里特医学院人工智能医学研究所)
  • Université Laval(拉瓦尔大学)
  • Mila – Quebec Artificial Intelligence Institute(Mila – 魁北克人工智能研究所)
  • Technische Universität Berlin(柏林工业大学)
  • Korea University(高丽大学)
  • Max Planck Institute for Informatics(马克斯·普朗克信息学研究所)

机构由 AI 辅助整理,请以论文原文为准。

↑