arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CADENCE:用于从心电图基础模型中提取可解释神经概念的心脏原子字典

CADENCE: A Cardiac Atom Dictionary for Interpretable Neural Concept Extraction from ECG Foundation Models

Yixuan Duan, Arjun Naik, Sadeer Al-Kindi, Wei Qiu

arXiv 2607.25244首次发表:更新:

发表机构

Rice University; Houston Methodist Hospital(莱斯大学; 休斯顿卫理公会医院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对12导联心电图基础模型生理知识不透明问题,提出CADENCE框架,用BatchTopK稀疏自动编码器分解模型,得到稀疏心脏原子,提升临床表型和形态学预测AUROC,能恢复生理关系、验证原子描述,为审查模型知识提供可扩展框架。

AI 中文摘要

12导联心电图(ECG)的基础模型在临床任务中表现良好,但其表征中编码的生理知识仍不透明。我们提出了CADENCE框架,将ECG基础模型分解为可人工解释、可查询的生理概念字典。利用BatchTopK稀疏自动编码器,CADENCE将来自900多万个ECG令牌的第6层嵌入分解为8192个稀疏心脏原子。这些原子在临床表型和波形形态方面比单个密集嵌入维度对齐得更好,能恢复心律失常、传导异常、梗死和复极模式、腔室和轴的发现以及特定导联和搏动阶段的波形基元。在第6层,最佳原子在临床表型上的平均AUROC为0.88,在形态学上为0.90,而最佳密集维度分别为0.78和0.83。稀疏原子探针在表型、形态学和年龄预测方面匹配或优于密集探针,同时将每个预测归因于一小部分可解释的原子;表型AUROC从0.93提高到0.95。原子空间几何恢复了生理上连贯的关系,有针对性的原子消融选择性地改变了冻结的下游输出。一个自动化的大语言模型管道通过预测保留的激活来生成并定量验证原子描述。在独立的外部ECG数据集上,CADENCE恢复了重叠概念并保持了一致的表型预测性能。CADENCE为发现和审查ECG基础模型编码的生理知识提供了一个可扩展的框架。

英文摘要

Foundation models for 12-lead electrocardiograms (ECGs) transfer well across clinical tasks, but the physiological knowledge encoded in their representations remains opaque. We present CADENCE, a framework that decomposes an ECG foundation model into a human-interpretable, queryable dictionary of physiological concepts. Using a BatchTopK sparse autoencoder, CADENCE factorizes Layer-6 embeddings from more than nine million ECG tokens into 8,192 sparse cardiac atoms. These atoms align better than individual dense embedding dimensions with clinical phenotypes and waveform morphology, recovering arrhythmias, conduction abnormalities, infarction and repolarization patterns, chamber and axis findings, and lead- and beat-phase-specific waveform primitives. At Layer 6, the best atoms achieve mean AUROCs of 0.88 for clinical phenotypes and 0.90 for morphology, versus 0.78 and 0.83 for the best dense dimensions. Sparse atom probes match or outperform dense probes for phenotype, morphology, and age prediction while attributing each prediction to a small set of interpretable atoms; phenotype AUROC improves from 0.93 to 0.95. Atom-space geometry recovers physiologically coherent relationships, and targeted atom ablation selectively changes frozen downstream outputs. An automated LLM pipeline generates and quantitatively validates atom descriptions by predicting held-out activations. On independent external ECG datasets, CADENCE recovers overlapping concepts and maintains consistent phenotype-prediction performance. CADENCE provides a scalable framework for discovering and auditing the physiological knowledge encoded by ECG foundation models.

Comments21 pages, 5 main figures, 15 appendix figures, 6 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑