arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.06285cs.LGcs.AIstat.ML

概念瓶颈模型解释何时是忠实且紧凑的?

When Are Concept Bottleneck Model Explanations Faithful and Compact?

Stefano Teso, Emanuele Marconato, Steve Azzolin, Antonio Vergari

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对概念瓶颈模型,论证忠实解释需包含全部概念,损害紧凑性,并提出概率建模与组套索稀疏化策略,以促进紧凑且忠实的解释,警示天真的可解释性主张。

中文摘要 AI 辅助

概念瓶颈模型(CBMs)是一种神经分类器,能够通过高层概念解释其决策,从而可能实现理解、引导和调试。然而,它们的解释通常是启发式推导的。基于形式化可解释性,我们认为这些解释也应当是忠实的,即不能错误地报告哪些概念真正起作用。我们证明,对于广泛使用的CBM架构,包括近期基于VLM的变体,忠实的解释必须包含瓶颈中的所有概念,当瓶颈规模较大时,这会损害可解释性。这一结果同时适用于启发式解释和通过构造保证忠实的正式解释。为了促进紧凑且忠实解释的存在,我们建议:i)将概念概率性地建模为二元或分类随机变量(而非logits),以及ii)在训练时通过组套索(group lasso)进行逐概念稀疏化(而非常规的弹性网络)。我们还扩展了形式化可解释性的算法到CBMs,并展示它们在保证和解释大小方面优于自然启发式方法。总体而言,我们的工作警示了天真的可解释性主张,并为确保CBMs如宣传的那样可解释提供了形式化条件和实用策略。

英文摘要

Concept bottleneck models (CBMs) are neural classifiers that allow to explain their decisions via high-level concepts, potentially enabling understanding, steering and debugging. However, their explanations are often derived heuristically. Building on formal explainability, we argue they should also be faithful, i.e., not misreport which concepts actually matter. We show that, for widespread CBM architectures, including recent VLM-based variants, faithful explanations must include all concepts in the bottleneck, compromising interpretability when this is large. This result applies to both heuristic and faithful-by-construction formal explanations. To encourage the existence of compact faithful explanations, we suggest i) modeling concepts probabilistically as binary or categorical random variables (rather than logits), and ii) employing per-concept training-time sparsification via group lasso (rather than regular elastic net). We also extend algorithms from formal explainability to CBMs, and show they outperform natural heuristics in terms of guarantees and explanation size. Overall, our work warns against naive interpretability claims and provides formal conditions and practical strategies for ensuring CBMs are as interpretable as advertised.

发表机构

  • University of Trento(特伦托大学)
  • University of Copenhagen(哥本哈根大学)
  • University of Edinburgh(爱丁堡大学)

机构由 AI 辅助整理,请以论文原文为准。

↑