arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越局部检查:基于指南的心电图分类事后可解释人工智能方法全局评估

Beyond Local Inspection: Global, Guideline-Grounded Evaluation of Post-hoc XAI Methods for ECG Classification

Nils Gumpfer, Michael Guckert, Samuel Sossalla, Birgit Aßmus, Jennifer Hannig

arXiv 2607.24035首次发表:更新:

发表机构

Hessian Center for Artificial Intelligence (hessian.AI); Technische Hochschule Mittelhessen, University of Applied Sciences; Justus-Liebig-University Giessen(黑森人工智能中心(hessian.AI); 中黑森应用科学大学; 吉森尤斯图斯-李比希大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对可解释人工智能在医学领域的问题,以心电图数据为对象,引入基于临床指南的全局框架评估事后XAI方法,通过实验揭示从计算机视觉转移来的方法有系统性失败,不同模式可靠性不一致,凸显全局评估能发现样本级热图不易察觉的解释问题。

AI 中文摘要

可解释人工智能(XAI)用于评估人工智能模型是否依赖有意义的模式,但对个体预测看似合理的解释可能会系统性地歪曲模型行为。在医学领域,这一问题尤为突出,模型可能依赖无关信号特征而非疾病特定模式且难以识别。我们利用心电图(ECG)数据应对这一挑战,临床指南为其提供了关于诊断相关信号区域的明确知识。我们引入了一个基于指南的全局框架,该框架汇总心跳间的解释,以对照临床定义的感兴趣区域对其进行评估。使用在PTB-XL上训练的四个二元分类器,我们评估了13种基于梯度的方法在两类模式上的表现:低振幅段和高振幅QRS形态。结果显示,从计算机视觉转移过来的方法存在系统性失败。它们的解释往往遵循信号幅度而非临床相关性,平均斯皮尔曼相关性高达0.69,导致它们忽略了具有诊断决定性的低振幅区域。对于缺血,LRP-ε仅将4.6%的相关性分配给ST段,而LRP-SIGN为63.8%。13种方法中有9种在至少一种情况下低于随机水平,表明不同模式间可靠性不一致。这些发现表明,基于领域的全局评估可以揭示样本级热图中不明显的系统性解释失败。

英文摘要

Explainable AI (XAI) is used to assess whether artificial intelligence models rely on meaningful patterns, yet explanations that appear plausible for individual predictions may systematically misrepresent model behavior. This is particularly problematic in medicine, where models may rely on irrelevant signal characteristics rather than disease-specific patterns without being recognizable. We address this challenge using electrocardiogram (ECG) data, for which clinical guidelines provide explicit knowledge about diagnostically relevant signal regions. We introduce a global, guideline-grounded framework that aggregates explanations across heartbeats to evaluate them against clinically defined regions of interest. Using four binary classifiers trained on PTB-XL, we assess 13 gradient-based methods across two categories of patterns: low-amplitude segments and high-amplitude QRS morphology. Our results reveal a systematic failure of methods transferred from computer vision. Their explanations often follow signal amplitude rather than clinical relevance, with mean Spearman correlations up to 0.69, leading them to overlook diagnostically decisive low-amplitude regions. For ischemia, LRP-$ε$ assigns only 4.6% of relevance to the ST segment, compared with 63.8% for LRP-SIGN. Nine of 13 methods fall below chance for at least one condition, indicating inconsistent reliability across patterns. These findings show that global, domain-grounded evaluation can uncover systematic explanation failures not obvious from sample-level heatmaps.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑