arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.16681cs.LG

一种通过自监督(JEPA)和联邦表示学习进行早期脓毒症预测的框架

A Framework for Early Sepsis Prediction via Self-Supervised (JEPA) and Federated Representation Learning

Umair bin Mansoor, Munaf Rashid, Roomi Naqvi

AI总结:

研究针对电子健康记录中早期脓毒症预测面临的问题,比较了JEPA等四种建模范式及原始特征基线,最佳模型在发病时接近SupMix基准且减少生物标志物使用,一级管道有显著提升,微调后的VICReg编码器特征持久性最佳。

AI中文摘要:

从电子健康记录中进行早期脓毒症预测面临不规则采样、高缺失率和类别不平衡等挑战。我们系统地比较了四种建模范式:通过掩码潜在预测的自监督联合嵌入预测架构(JEPA)、具有双视图增强的自监督VICReg(方差-不变性-协方差正则化)、VICReg预训练编码器的半监督微调以及监督时间卷积网络(TCN),并与原始特征基线进行比较。所有模型都共享一个共同的预处理管道,对从MIMIC-III数据集中通过稀疏性分析选择的7种生物标志物进行每小时分箱并应用前向填充插补。我们的最佳模型(JEPA + XGBoost + 均值池化)在发病时(H0)达到AUPRC 0.636,接近SupMix基准(0.667),同时使用的生物标志物减少了83%。一级管道(VICReg预训练,然后进行半监督微调并使用XGBoost)在H时达到AUPRC 0.510,比原始特征基线(0.165)提高了3.1倍,比端到端监督TCN(0.474)提高了7.6%。至关重要的是,微调后的VICReg编码器表现出最具时间持久性的表示,从H0到H10仅下降16.8%,而监督TCN为47.5%,JEPA为65.3%,这表明具有任务感知微调的自监督预训练产生的特征在发病时既尖锐又在整个预测范围内稳健。

英文摘要:

Early sepsis prediction from electronic health records is challenged by irregular sampling, high missingness, and class imbalance. We systematically compare four modeling paradigms -- self-supervised Joint Embedding Predictive Architecture (JEPA) via masked latent prediction, self-supervised VICReg (variance-invariance-covariance regularization) with two-view augmentation, semi-supervised fine-tuning of a VICReg-pretrained encoder, and supervised Temporal Convolutional Network (TCN) -- alongside raw-feature baselines. All models share a common preprocessing pipeline of hourly binning with forward-fill imputation applied to 7 biomarkers selected via sparsity analysis from the MIMIC-III dataset. Our best model (JEPA + XGBoost + mean pooling) achieves AUPRC 0.636 at the time of onset (H0), approaching the SupMix benchmark (0.667) while using 83\% fewer biomarkers. The Tier 1 pipeline -- VICReg pretraining followed by semi-supervised fine-tuning and XGBoost -- achieves AUPRC 0.510 at H0, a 3.1$\times$ improvement over the raw-feature baseline (0.165) and a 7.6\% improvement over the end-to-end supervised TCN (0.474). Crucially, the fine-tuned VICReg encoder exhibits the most temporally persistent representations, degrading only 16.8\% from H0 to H10 compared to 47.5\% for supervised TCN and 65.3\% for JEPA, demonstrating that self-supervised pretraining with task-aware fine-tuning yields features that are both sharp near onset and robust across prediction horizons.

补充信息

↑