arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38924cs.CV

从图像解读到临床推理:基于上游医生情境感知的多模态学习与因果强化学习

From Image Interpretation to Clinical Reasoning: Upstream Physician-Context-Aware Multimodal Learning with Causal Reinforcement Learning

Jialu Pi, Yanan Ma, Weijie Chen, Owen Crystal, Shubham Trivedi, Stephen Xie, Anna Silverman, Matthew Stib, Chadi Ayoub, Reza Arsanjani, Imon Banerjee

首次发表
浏览论文内容

中文总结 AI 辅助

提出一种因果强化学习框架,整合胸部X光片和临床病史进行机会性MACE预测,通过角色解耦双LLM架构和双动作策略,在多个数据集上优于现有模型,显著提升推理质量。

中文摘要 AI 辅助

主要不良心血管事件(MACE)仍是全球范围内导致死亡的主要原因。利用常规采集的临床数据进行机会性筛查,为在急性事件发生前识别高危个体提供了一种可扩展的方法。尽管胸部X光片(CXR)能够捕捉潜在的心血管生物标志物,且临床病史提供了互补的患者背景信息,但现有的医学视觉语言模型主要针对放射学解读而非预后推理进行优化。我们提出了一种用于多模态临床推理的因果强化学习框架,该框架整合了CXR和医生撰写的临床病史,用于机会性MACE预测。该框架引入了(1)一种角色解耦的双大语言模型(LLM)架构,将推理与风险预测分离;(2)一种双动作因果强化学习策略,用于证据选择和推理优化;以及(3)因果标记剪枝,以学习紧凑的多模态表示。在内部队列、急诊科队列以及外部MIMIC数据集上进行的评估中,所提出的框架 consistently 优于单模态基线和最先进的医学视觉语言模型,分别实现了0.720、0.760和0.845的AUROC。它还显著提高了推理质量,获得了更高的GREEN分数和更高的专家偏好,同时在多样化的患者群体中保持了稳健的预测性能。

英文摘要

Major adverse cardiovascular events (MACE) remain the leading cause of mortality worldwide. Opportunistic screening using routinely acquired clinical data offers a scalable approach for identifying high-risk individuals before acute events occur. Although chest X-rays (CXRs) capture latent cardiovascular biomarkers and clinical histories provide complementary patient context, existing medical vision-language models are primarily optimized for radiology interpretation rather than prognostic reasoning. We propose a causal reinforcement learning framework for multimodal clinical reasoning that integrates CXRs and physician-authored clinical histories for opportunistic MACE prediction. The framework introduces (1) a role-decoupled dual-LLM architecture that separates reasoning from risk prediction, (2) a dual-action causal reinforcement learning policy for evidence selection and reasoning optimization, and (3) causal token pruning to learn compact multimodal representations. Evaluated on an internal cohort, an emergency department cohort, and the external MIMIC dataset, the proposed framework consistently outperformed unimodal baselines and state-of-the-art medical vision-language models, achieving AUROCs of 0.720, 0.760, and 0.845, respectively. It also substantially improved reasoning quality, achieving higher GREEN scores and higher expert preference while maintaining robust predictive performance across diverse patient populations.

发表机构

  • Mayo Clinic(梅奥诊所)
  • Arizona State University(亚利桑那州立大学)

机构由 AI 辅助整理,请以论文原文为准。

↑