arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一种用于12导联心电图分类的混合CNN-状态空间-注意力骨干网络与联合嵌入预测预训练

A Hybrid CNN--State-Space--Attention Backbone with Joint-Embedding Predictive Pretraining for 12-Lead ECG Classification

Yakoub Bazi, Sarah Aljuhani, Mohamad M. Al Rahhal, Mansour Zuair, Naif Alajlan

arXiv 2609.29376首次发表:更新:

发表机构

King Saud University(沙特国王大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出一种混合CNN-状态空间-注意力骨干网络,结合面向心电图的JEPA预训练,在紧凑参数下实现12导联ECG分类,并在多个数据集上提升迁移性能。

AI 中文摘要

自动12导联心电图(ECG)分类需要能够同时捕捉局部波形形态、长程时间动态和跨导联依赖关系的表示,然而在单一高效架构中整合这些特性仍然具有挑战性。本文提出了一种用于12导联心电图分类的混合CNN-SSM-Attention骨干网络。卷积主干执行早期波形令牌化和时间缩减,混合状态空间和深度卷积块建模时间动态和局部形态,后期自注意力阶段在降低的分辨率下实现全局令牌交互。为了改善从无标签数据的迁移,我们进一步开发了一种面向心电图的联合嵌入预测预训练(JEPA)框架。与基于ViT的JEPA方法在编码器之前掩蔽补丁令牌不同,所提出的方法在潜在时间分辨率上采样跨度掩码,并将其投影回波形域,然后从动量编码器预测干净的潜在目标,而无需波形重建。在CPSC2018、Chapman-Shaoxing和PTB-XL上的实验,使用约35万条无标签的CODE-15记录进行预训练,表明所提出的骨干网络在紧凑参数预算下提供了强大的监督基线。JEPA预训练进一步改善了迁移性能,特别是在减少标签的设置中,以及在完全微调和基于LoRA的适应下均如此。代码:此https URL

英文摘要

Automatic 12-lead electrocardiogram (ECG) classification requires representations that jointly capture local waveform morphology, long-range temporal dynamics, and cross-lead dependencies, yet integrating these properties within a single efficient architecture remains challenging. This paper introduces a hybrid CNN-SSM-Attention backbone for 12-lead ECG classification. A convolutional stem performs early waveform tokenization and temporal reduction, mixed state-space and depthwise-convolutional blocks model temporal dynamics and local morphology, and a late self-attention stage enables global token interaction at reduced resolution. To improve transfer from unlabeled data, we further develop an ECG-oriented Joint-Embedding Predictive Pretraining (JEPA) framework. Unlike ViT-based JEPA methods that mask patch tokens before the encoder, the proposed method samples span masks at the latent temporal resolution and projects them back to the waveform domain, then predicts clean latent targets from a momentum encoder without waveform reconstruction. Experiments on CPSC2018, Chapman-Shaoxing, and PTB-XL, with pretraining on approximately 350K unlabeled CODE-15 recordings, show that the proposed backbone provides strong supervised baselines under a compact parameter budget. JEPA pretraining further improves transfer, particularly in reduced-label settings and under both full fine-tuning and LoRA-based adaptation. Code: https://github.com/yakoubbazi/Hybrid_ECG_Jepa

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑