用于临床推理和证据生成的放射影像世界模型
A radiographic world model for clinical reasoning and evidence generation
- The University of Chicago(芝加哥大学)
- Emory University(埃默里大学)
- Memorial Sloan Kettering Cancer Center(纪念斯隆-凯特琳癌症中心)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出放射影像世界模型MedDream,通过共享潜在状态统一诊断与生成,在八数据集上优于现有方法,并实现基于亚组差距的定向证据生成,提升临床性能。
AI中文摘要:
医学影像人工智能(AI)通常被开发为从放射影像到诊断输出的独立映射,或从临床描述到生成影像的独立映射,尽管两者源于相同的底层放射影像状态。世界模型公式转而寻求学习这一状态的内部表征,以支持临床读出和放射影像观察的条件模拟。在此,我们介绍了MedDream,一种放射影像世界模型,它从配对的胸部放射影像-文本观察中学习共享的连续潜在状态,用于诊断推理和报告条件下的证据生成。MedDream在从440万候选中筛选出的265万对无泄漏胸部放射影像-文本对上进行了预训练。在八个临床数据集和两个独立读者队列中,MedDream优于领先的诊断和生成比较器。在诊断推理方面,MedDream在疾病识别、标签稀缺适应、严重程度评估和定位方面表现出强大的泛化能力,而MedDream支持的审查将住院医师与独立放射科医生共识的平均一致性从56.3%提高到63.0%。在证据生成方面,MedDream生成的放射影像保留了临床相关的病理特征,并提高了在保留的真实数据上的下游性能,合成增强将外部VinDr-CXR宏AUROC从76.4%提高到81.4%。更重要的是,基于预设亚组性能差距的条件生成实现了有针对性的证据构建,将亚洲患者的加权F1提高了3.1个百分点,而匹配体积的无引导增强则使其降低了2.3个百分点。这些发现确立了放射影像世界模型作为医学AI的一条路径,使医学AI学习有临床意义的内部状态,用于解释、模拟和构建临床使用的证据。
英文摘要:
Medical imaging artificial intelligence (AI) is commonly developed as separate mappings from radiographs to diagnostic outputs or from clinical descriptions to generated images, although both arise from the same underlying radiographic state. A world-model formulation instead seeks to learn an internal representation of this state that can support both clinical readout and conditional simulation of radiographic observations. Here we introduce MedDream, a radiographic world model that learns a shared continuous latent state from paired chest radiograph-text observations for diagnostic reasoning and report-conditioned evidence generation. MedDream was pretrained on 2.65 million leakage-controlled chest radiograph-text pairs curated from 4.40 million candidates. Across eight clinical datasets and two independent reader cohorts, MedDream outperformed leading diagnostic and generative comparators. For diagnostic reasoning, MedDream showed strong generalization across disease recognition, label-scarce adaptation, severity assessment, and localization, while MedDream-supported review increased mean resident concordance with independent radiologist consensus from 56.3% to 63.0%. For evidence generation, MedDream produced radiographs that preserved clinically relevant pathology and improved downstream performance on held-out real data, with synthetic augmentation increasing external VinDr-CXR macro-AUROC from 76.4% to 81.4%. More importantly, conditioning generation on prespecified subgroup performance gaps enabled targeted evidence construction, increasing weighted F1 by 3.1 percentage points in Asian patients, whereas matched-volume unguided augmentation decreased it by 2.3 points. These findings establish radiographic world models as a path toward medical AI that learns clinically meaningful internal states for interpreting, simulating, and constructing evidence for clinical use.