arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38625cs.LGcs.AI

可解释但脆弱?几何-语义扰动下概念瓶颈模型的鲁棒性

Interpretable but Fragile? Robustness of Concept Bottlenecks under Geometric-Semantic Perturbations

Hanwei Zhang, Tianma Hu, Gaojie Jin, Xu Cheng, Ronghui Mu

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过生成器框架比较概念瓶颈模型与标准分类器在几何和语义扰动下的鲁棒性,揭示可解释性不固有提升鲁棒性,而是重新分配敏感性,存在依赖扰动机制的任务结构权衡。

中文摘要 AI 辅助

概念瓶颈模型(CBMs)旨在提供可解释的中间表示,然而这种瓶颈如何影响鲁棒性仍不清楚,现有研究报道了混合且有时相互矛盾的发现。我们认为这些差异源于混淆了不同的鲁棒性概念和扰动机制,而非对CBMs本身的根本性分歧。为厘清这些因素,我们引入了一个基于生成器的评估框架,该框架能够在两种不同的扰动类型下对标准分类器和CBMs进行受控比较:潜在空间中的连续几何扰动和概念空间中的离散语义干预。在此框架内,我们通过预测和概念级敏感性指标进行经验评估,并利用潜在空间和概念空间中的随机平滑进行可认证评估。通过实验,我们澄清了先前相互矛盾的发现,阐明了概念瓶颈在何时以及何种意义上改善或不改善鲁棒性。通过进一步分析不同任务条件下的鲁棒性,包括类别语义相似性和概念词汇表大小,我们表明可解释性并不固有地赋予鲁棒性。相反,概念瓶颈转移了敏感性表现的位置和方式,揭示了一种微妙的可解释性-鲁棒性权衡,该权衡关键取决于扰动机制和任务结构。综合来看,我们的结果表明可解释性和鲁棒性是不同目标:可解释的中间表示并不均匀地改善鲁棒性,而是将敏感性重新分布在扰动空间和模型族中。

英文摘要

Concept Bottleneck Models (CBMs) are designed to provide interpretable intermediate representations, yet how such bottlenecks affect robustness remains unclear, with existing studies reporting mixed and sometimes contradictory findings. We argue that these discrepancies arise from conflating different robustness notions and perturbation regimes, rather than from fundamental disagreements about CBMs themselves. To disentangle these factors, we introduce a generator-based evaluation framework that enables controlled comparisons between standard classifiers and CBMs under two distinct perturbation types: continuous geometric perturbations in latent space and discrete semantic interventions in concept space. Within this framework, we evaluate robustness both empirically, via prediction and concept-level sensitivity metrics, and certifiably, using randomized smoothing in latent and concept spaces. Across experiments, we reconcile previously conflicting findings by clarifying when, and in what sense, concept bottlenecks do or do not improve robustness. By further analyzing robustness under varying task conditions, including class semantic similarity and concept vocabulary size, we show that interpretability does not inherently confer robustness. Instead, concept bottlenecks shift where and how sensitivity manifests, revealing a nuanced interpretability robustness trade off that depends critically on the perturbation regime and task structure. Together, our results show that interpretability and robustness are distinct objectives: interpretable intermediate representations do not uniformly improve robustness, but instead redistribute sensitivity across perturbation spaces and model families.

发表机构

  • Saarland University(萨尔兰大学)
  • Tianjin University of Technology(天津工业大学)
  • University of Macau(澳门大学)
  • University of Exeter(埃克塞特大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑