arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一种用于多变量时间序列的轻量级自监督学习框架:基于层次JEPA的心电图数据方法

Hierarchical Self-Supervised Representation Learning Framework for Multivariate Time Series Grounded in ECG Analysis

Siwon Kim

arXiv 2607.01145首次发表:更新:

发表机构

Research Institute of Basic Sciences, Seoul National University(首尔大学基础科学研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出ER-JEPA,一种轻量级自监督学习框架,通过层次化联合嵌入预测架构对多变量时间序列进行表征学习,在心电图数据上实现高效预训练和下游任务最优性能。

AI 中文摘要

医学领域的数据分析常面临目标数据集有限而大量未标注数据分布广泛的情况。在这种情况下,自监督学习方法能有效利用大数据集,成为心电图分析的热门选择。本文提出事件重建联合嵌入预测架构(ER-JEPA),一种用于多变量时间序列的轻量级自监督学习框架,其名称和双层层次结构受心脏病专家诊断方法的启发。ER-JEPA的核心包括:(1)两阶段结构,先为每个时间间隔构建表征,再将这些表征作为单变量时间序列处理;(2)两个联合嵌入预测架构的层次化集成;(3)视觉Transformer骨干网络。两个JEPA的结构串联将模型归类为层次JEPA(H-JEPA),旨在编码多级抽象表征以增强复杂任务的预测能力。本研究报告了H-JEPA在12导联心电图数据(作为多变量时间序列)上的成功应用,并分析了预训练阶段层次表征的敏感性。该模型在约18万条10秒记录上预训练,在ST-MEM基准测试中实现了最先进的下游性能,且计算速度快、资源占用极少。

英文摘要

Data analysis in the medical domain often encounters scenarios involving a limited target dataset and a large, unannotated dataset with a general distribution. Under such circumstances, self-supervised learning (SSL) methods are highly effective for utilizing large datasets, making them a popular choice for electrocardiogram (ECG) analysis. This work presents the Event Reconstruction Joint-Embedding Predictive Architecture (ER-JEPA), a lightweight SSL framework for multivariate time series, whose name and two-fold hierarchical structure are inspired by the diagnostic approach of cardiologists. At its core, ER-JEPA features: (1) a two-stage structure that constructs representations for each time interval and subsequently processes these representations as a univariate time series, (2) the hierarchical integration of two Joint-Embedding Predictive Architectures (JEPAs), and (3) a Vision Transformer (ViT) backbone. The structural concatenation of two JEPAs categorizes the model as a Hierarchical JEPA (H-JEPA), designed to encode multiple levels of abstract representations for enhanced prediction on complex tasks. This study reports a successful application of H-JEPA to 12-lead ECG data as a multivariate time series, alongside an analysis of the sensitivity of hierarchical representation during the pretraining stage. Furthermore, this study provides a qualitative demonstration that the intermediate representations produced by the first module of ER-JEPA excel at local feature extraction, as they are structurally free from over-smoothing. Pretrained on approximately 180,000 10-second recordings, the model achieves state-of-the-art downstream performance on the ST-MEM benchmark, with rapid computation and minimal resource usage.

Comments29 pages, 8 figures. Further clarified, added the new downstream task in the abstract and intro since last update.<< Polished text, improved formatting, fixed speed benchmark result, and added new downstream task. Code will be made publicly available soon

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑