arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于卷积神经网络的心电图图像分类中的捷径学习与聪明汉斯效应分析

Analysis of the Shortcut Learning and Clever Hans Effect in CNN based ECG Image Classification

Abhay Kumar Pathak, Mrityunjay Chaubey, Manjari Gupta, Deepti Mishra

arXiv 2607.25117首次发表:更新:

发表机构

Banaras Hindu University; Norwegian University of Science and Technology (NTNU); University of Petroleum and Energy Studies(贝拿勒斯印度教大学; 挪威科技大学; 石油与能源研究大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究基于卷积神经网络,在公开心电图图像数据集上,通过创建六个图像衍生特征集,计算相关分数和测试结果,评估心电图图像分类器是学习临床有意义形态还是捷径线索,以分析捷径学习和聪明汉斯效应。

AI 中文摘要

用于心电图图像分类的深度学习模型可能通过利用非生理视觉线索而非心电图波形形态来实现高精度。鉴于深度学习模型的黑箱性质,其高预测性能往往难以转化为临床或现实世界中的信任、可解释性和可操作的决策。本研究使用卷积神经网络在一个公开可用的心电图图像数据集上研究捷径学习和聪明汉斯效应。过程中创建了六个图像衍生特征集,用于测试去除波形信息或引入人工特定类伪影时分类性能是否持续。计算了跨特征集表示的捷径保留分数、预测一致性和置信度差异,以评估学习模式的透明度。还给出平均积分梯度和遮挡敏感度测试结果以检查模型归因是否聚焦于心电图相关波形区域或非临床伪影。通过特征集的性能变化和归因模式来识别潜在的聪明汉斯行为。本研究评估心电图图像分类器是学习临床有意义的形态还是由报告布局、元数据、对比度、模糊或人工标记引入的捷径线索。

英文摘要

Deep learning models for ECG image classification may achieve high accuracy by exploiting non-physiological visual cues instead of ECG waveform morphology. Given the black-box nature of deep learning models, their promise of high predictive performance often remains insufficiently translated into clinical or real-world trust, interpretability, and actionable decision-making. In this study, we examine shortcut learning and Clever Hans effect in a publicly available ECG image dataset using convolutional neural networks. In process we have created six image-derived feature sets (FSs), FS1: raw full ECG images, FS2: cropped waveform-only images, FS3: waveform-masked metadata images, FS4: red-arrow artifact images for the myocardial infarction class, FS5: contrast-enhanced images for the abnormal heartbeat class and FS6: Gaussian-blurred images for the normal class. These controlled representations were used to test whether classification performance persists when waveform information is removed or when artificial class-specific artifacts are introduced. Shortcut retention score, prediction consistency and confidence divergence across Feature-Set Representations have been calculated to assess the transparency about the learning pattern. Along with factual results, average Integrated Gradients and occlusion sensitivity test results are presented to inspect whether model attribution focused on ECG-relevant waveform regions or on non-clinical artifacts. Performance changes across feature sets and attribution patterns were used to identify potential Clever Hans behavior. This study evaluates whether ECG image classifiers learn clinically meaningful morphology or shortcut cues introduced by report layout, metadata, contrast, blur, or artificial markers.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑