阅读整个心脏:用于多模态心脏表征学习的潜在注意力掩码自编码器
Reading the Whole Heart: Latent-Attention Masked Autoencoders for Multimodal Cardiac Representation Learning
浏览论文内容
中文总结 AI 辅助
提出潜在注意力掩码自编码器(LAMAE),通过潜在空间跨模态信息交换和结构感知预训练,在120万住院数据上实现多模态心脏表征学习,优于现有基线并稳健处理缺失模态。
中文摘要 AI 辅助
心血管诊断依赖于整合互补模态,如心电图、超声心动图、胸部X光片和临床变量,每种模态捕捉心脏生理学中不同但相关的方面。然而,大多数医学基础模型仍然局限于特定模态,仅在微调或训练后阶段才组合模态。这丢弃了临床医生自然整合的跨模态证据,并忽略了每种模态内部的结构。我们引入了潜在注意力掩码自编码器(LAMAE),一种多模态、结构感知的掩码自编码器,在自监督预训练期间联合学习患者级别的表征。LAMAE并非事后融合模态,而是通过一个共享的潜在注意力模块,在潜在空间中直接交换信息,该模块在“研究-视图-实体”层级上运行,从而能够聚合可变观测并优雅地处理缺失模态。在超过120万次MIMIC-IV医院住院数据上预训练后,LAMAE在多模态医院住院任务(如院内死亡率、ICD-10和DRG编码以及住院时长)上优于特定模态的预训练和强对比学习及视觉-语言基线,同时在单模态任务上保持竞争力。即使在测试时仅有一种模态可用,这些优势依然存在,表明对模态内和模态间结构进行建模能够产生更稳健、更具可迁移性的表征。
英文摘要
Cardiovascular diagnosis and treatment rest on integrating complementary modalities, such as electrocardiogram, echocardiography, and chest X-rays, each capturing distinct but complementary aspects of cardiac pathophysiology. Yet most medical foundation models remain modality-specific, combining modalities only for finetuning or post-training. This discards the cross-modal evidence clinicians naturally integrate and ignores the structure within each modality. We introduce Latent-Attention Masked Autoencoder (LAMAE), a multimodal, structure-aware masked autoencoder that jointly learns patient-level representations during self-supervised pretraining. Instead of fusing modalities post hoc, LAMAE exchanges information directly in the latent space through a shared latent-attention module operating over a study-view-entity hierarchy, enabling aggregation of variable observations and handling of missing modalities. Pretrained on over 500'000 MIMIC-IV hospital stays, LAMAE outperforms modality-specific pretraining and strong contrastive and vision-language baselines across multimodal hospital-stay tasks, such as in-hospital mortality, ICD-10 and DRG coding, and length of stay, while remaining competitive on unimodal tasks. Modeling this structure also pays off within a single modality: even without cross-modal information, the latent-attention module improves representations over modality-specific pretraining.
发表机构
- ETH Zurich(苏黎世联邦理工学院)
- University of Basel(巴塞尔大学)
机构由 AI 辅助整理,请以论文原文为准。