arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.05202eess.IVcs.CV

真实世界多模态纵向肺癌数据集

Real-World Multi-Modal and Longitudinal Lung Cancer Dataset

Rita Cordeiro Mendes, Maria Rita Fonseca Verdelho, Carlos Santiago, Catarina Barata

首次发表
浏览论文内容

中文总结 AI 辅助

本研究发布含1365名肺癌患者的多中心多模态纵向数据集,提供相关基准,证实多模态融合可提升缺失数据下的肺癌生存预测性能。

中文摘要 AI 辅助

多模态学习通过整合医学影像、临床记录、基因组学等异构数据源,在医疗应用中展现出强大潜力,可提升预测性能并支持临床决策。然而该领域的进展常受两大关键挑战制约:一是经过精心整理、可直接使用且能准确反映真实世界状况的数据集有限,真实世界中的医疗数据常存在收集不一致、不完整的问题;二是异构数据模态的固有整合难度。本研究中,我们引入了一个新整理的多中心、多模态、纵向数据集,旨在支持在真实条件下对各类学习流程进行评估。该数据集共包含1365名肺癌患者,涵盖三种影像模态(全切片图像、CT扫描、PET扫描)、结构化临床数据、转录组数据,以及纵向随访和治疗信息。每种影像模态对应数据集包含不止一个实例,且该数据集在各模态间存在大量且非均匀的缺失值,非常适合用于研究鲁棒的多模态融合策略。我们还提供了单模态和多模态基准,用于12个月总生存期预测、疾病特异性生存期预测,以及在严重缺失数据下的风险预测纵向基准。结果显示,尽管缺失值水平较高,整合互补模态始终比单模态方法提升了预测性能,凸显了真实临床环境下多模态融合的价值。该数据集和基准代码可在此处的URL获取。

英文摘要

Multi-modal learning has demonstrated strong potential in medical applications by integrating heterogeneous data sources such as medical imaging, clinical records, and genomics to improve predictive performance and support clinical decision-making. However, advances in this area are often constrained by two key challenges: the limited availability of well-curated, ready-to-use datasets that accurately reflect real-world conditions, where medical data are frequently collected inconsistently and are often incomplete; and the inherent difficulty of integrating heterogeneous data modalities. In this work, we introduce a newly curated multi-center, multi-modal, and longitudinal dataset designed to support the evaluation of a wide range of learning pipelines under realistic conditions. The dataset comprises a total of 1,365 lung cancer patients and has three imaging modalities (whole-slide images, CT scans, and PET scans), structured clinical data, transcriptomic, and longitudinal follow-up and treatment information. For each imaging modality the dataset contains more than one instance. Moreover, the dataset exhibits substantial and non-uniform missingness across modalities, making it well-suited for studying robust multi-modal fusion strategies. We further provide both uni-modal and multi-modal benchmarks on the task of 12-month overall survival prediction, disease-specific survival, as well as longitudinal benchmark of hazard prediction under severe missing data. Our results show that, despite high levels of missingness, integrating complementary modalities consistently improves predictive performance over uni-modal approaches, highlighting the value of multi-modal fusion in realistic clinical settings. The dataset and benchmark code are available at https://github.com/ritacmendes/MMIST-LUNG.

发表机构

  • Institute for Systems and Robotics, Instituto Superior Técnico(系统与机器人研究所,里斯本高等技术学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑