AI 中文总结
本文针对现有医学视觉-语言任务无法兼顾医学事实性与患者语境沟通的问题,提出PMRI任务及G-CARL框架,构建MMedReport基准,实验证实其解读更贴合患者需求。
AI 中文摘要
患者对医学报告的个性化解读需求日益凸显,满足该需求既需要基于证据的医学事实性,又需要结合语境的患者沟通,而现有的医学视觉-语言任务未能充分捕捉这两项核心要求。为弥合这一差距,本文提出了面向患者的医学报告解读(Patient-oriented Medical Report Interpretation,PMRI)这一新型开放式多模态生成任务,要求模型基于用户查询和对话历史,以准确且易懂的语言解读医学报告。这两个目标在可验证性上存在根本差异,但又紧密关联,导致在传统的监督微调及整体强化学习范式下难以联合优化。为应对这一挑战,本文提出了G-CARL,一种基于 grounded 核对清单对齐的强化学习框架,该框架将多源检索用于原子声明验证,结合上下文感知的实例特定加权核对清单以确保响应覆盖范围,为事实性、用户需求满足度和表达质量提供结构化监督,同时不限制响应多样性。本文还构建了真实世界的PMRI基准MMedReport,以及由临床医生设计的三维评估协议。大量实验表明,G-CARL在整体质量、声明级精度和核对清单召回率上始终优于现有的后训练基线,临床医生的成对偏好评估进一步证实,G-CARL生成的解读更准确,且更贴合患者需求。
英文摘要
Personalized interpretation of medical reports has emerged as an increasingly important need among patients. Addressing this need requires both evidence-grounded medical factuality and context-dependent patient communication, yet existing medical vision-language tasks do not adequately capture these dual requirements. To bridge this gap, we introduce Patient-oriented Medical Report Interpretation (PMRI), a novel open-ended multimodal generation task that requires models to explain medical reports in accurate and accessible language based on a user's query and dialogue history. These two objectives differ fundamentally in their verifiability, yet remain tightly coupled, making them difficult to optimize jointly under conventional supervised fine-tuning and holistic reinforcement learning paradigms. To address this challenge, we propose G-CARL, a grounded, checklist-aligned reinforcement learning framework that combines multi-source retrieval for atomic claim verification with context-aware, instance-specific weighted checklists for response coverage, providing structured supervision for factuality, user-demand satisfaction, and expression quality without constraining response diversity. We further construct MMedReport, a real-world PMRI benchmark, along with a clinician-designed three-dimensional evaluation protocol. Extensive experiments demonstrate that G-CARL consistently outperforms existing post-training baselines in overall quality, claim-level precision, and checklist recall. Pairwise preference evaluation by clinicians further confirms that G-CARL produces interpretations that are more accurate and better aligned with patient needs.