arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.27404cs.LGcs.AI

ECG-InterpBench:采用匹配规模稀疏自编码器评估心电图基础模型的可解释性

ECG-InterpBench: Benchmarking the Interpretability of ECG Foundation Models with Matched-Scale Sparse Autoencoders

发表机构莱斯大学
查看机构详情
  • Rice University(莱斯大学)

机构由 AI 辅助整理,请以论文原文为准。

Yixuan Duan, Wei Qiu

首次发表
浏览论文内容

中文总结 AI 辅助

ECG-InterpBench是评估心电图基础模型可解释性的基准,采用匹配规模稀疏自编码器,在多维度对比中揭示模型可解释性差异,补充了性能导向的心电图基准。

中文摘要 AI 辅助

现有的心电图基础模型基准主要评估下游预测性能,难以深入洞察其内部表示是否能被忠实分解、临床解读或在独立分析中复现。本文提出ECG-InterpBench,这一基准旨在系统评估心电图基础模型表示的可解释性。ECG-InterpBench将稀疏自编码器作为标准化测量工具,并使各模型的稀疏自编码器容量匹配,以实现受控对比。我们在5种标准化编码器深度、5种匹配字典宽度和3种随机种子下评估6种冻结的心电图基础模型,生成包含75个完全匹配的六模型对比块、共450个单元的可解释性图谱。该基准评估表示可解释性的互补维度,包括稀疏重建保真度、单特征可访问性与49种临床有意义心电图测量指标的覆盖度,以及跨种子特征可复现性。评估还量化了患者采样不确定性、深度与种子依赖的变化,以及对稀疏性参数化的敏感性。该基准显示,心电图基础模型呈现出不同的可解释性特征。在MIMIC-IV-ECG上的匹配复现确认,重建保真度与临床可访问性识别出不同的领先模型。该基准附带可执行评估代码、标准化清单、单元级指标及可复现性审计,ECG-InterpBench补充了以性能为中心的心电图基准,提供了容量受控且可复现的框架,用于在表示可解释性的不同维度对比心电图基础模型。

英文摘要

Existing benchmarks for electrocardiogram foundation models primarily evaluate downstream predictive performance, providing limited insight into whether their internal representations can be faithfully decomposed, clinically interpreted, or reproduced across independent analyses. We introduce ECG-InterpBench, a benchmark designed to systematically evaluate the interpretability of ECG foundation-model representations. ECG-InterpBench uses sparse autoencoders as standardized measurement instruments and matches their capacity across models to enable controlled comparisons. We evaluate six frozen ECG foundation models across five standardized encoder depths, five matched dictionary widths, and three random seeds, producing a 450-cell interpretability atlas comprising 75 exactly matched six-model comparison blocks. The benchmark evaluates complementary dimensions of representation interpretability, including sparse reconstruction fidelity, single-feature accessibility and coverage of 49 clinically meaningful ECG measurements, and cross-seed feature reproducibility. The evaluation further quantifies patient-sampling uncertainty, depth- and seed-dependent variation, and sensitivity to the sparsity parameterization. The benchmark reveals that ECG foundation models exhibit distinct interpretability profiles. A matched replication on MIMIC-IV-ECG confirms that reconstruction fidelity and clinical accessibility identify different leading models. The benchmark is accompanied by executable evaluation code, standardized manifests, cell-level metrics, and reproducibility audits. ECG-InterpBench complements performance-centered ECG benchmarks by providing a capacity-controlled and reproducible framework for comparing ECG foundation models across distinct dimensions of representation interpretability.

补充信息

↑