发表机构
Huazhong University of Science and Technology(华中科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有医疗报告生成方法的临床事实错误与视觉-语言对齐不足问题,提出DPO-Clin框架,通过ECD模块、M²DPO及反事实偏好数据构建提升模型可靠性,在多医学数据集上取得优异性能。
AI 中文摘要
尽管医疗报告生成(MRG)已取得显著进展,但事实错误的普遍存在仍限制了其可靠性。直接偏好优化(DPO)作为一种有前景的后训练范式,可提升监督微调(SFT)MRG模型的性能,然而现有基于DPO的MRG方法通常采用简单的偏好构建策略,直接将模型生成的报告与真实报告配对。该策略无意间将关键临床发现与临床无关的语言特征纠缠在一起,且根本缺乏显式的视觉-语言对齐。为解决这些挑战,我们提出DPO-Clin,一种新型后训练框架,其将偏好优化聚焦于临床发现与跨模态对齐。首先,我们引入实体级临床诊断(ECD)模块,执行精确的实体级事实诊断,ECD指导生成语言对齐的报告偏好对,将临床差异与语言变异隔离开。其次,为实现细粒度跨模态对齐,我们开发M²DPO,一种检索增强的多模态DPO变体,其强制由视觉上下文切换触发的文本偏好反转。第三,我们定位正确但高度不确定的预测实体,并应用反事实修改构建针对性偏好数据以缓解潜在风险,从而进一步提升模型可靠性。在两个公开胸部X射线数据集(MIMIC-CXR和IU X-Ray)及一个内部内窥镜数据集上开展的大量实验表明,DPO-Clin在临床感知指标上显著优于SFT基线,且相比现有基于DPO的MRG方法实现了更优性能,在不同基线架构及多样医学成像模态中展现出稳健的泛化能力。
英文摘要
Despite significant advances in Medical Report Generation (MRG), the reliability remains constrained by the prevalence of factual errors. While Direct Preference Optimization (DPO) has emerged as a promising post-training paradigm to enhance the performance of Supervised Fine-Tuned (SFT) MRG models, existing DPO-based MRG methods typically adopt a naive preference construction that directly pairs model-generated reports with ground truth reports. This strategy inadvertently entangles critical clinical findings with clinically irrelevant linguistic characteristics, and fundamentally lacks explicit vision-language alignment. To address these challenges, we propose DPO-Clin, a novel post-training framework that focuses preference optimization on clinical findings and cross-modal alignment. First, we introduce the Entity-level Clinical Diagnostic (ECD) module to perform a precise entity-level factual diagnosis. ECD guides the generation of linguistically-aligned report preference pairs, isolating clinical discrepancies from linguistic variations. Second, to achieve fine-grained cross-modal alignment, we develop M2DPO, a retrieval-augmented multi-modal DPO variant that enforces textual preference inversion triggered by visual context switches. Third, we locate correct yet highly uncertain predicted entities and apply counterfactual modifications to construct targeted preference data for latent risk mitigation, thereby further enhancing the model reliability. Extensive experiments on two public chest X-ray datasets (MIMIC-CXR and IU X-Ray) and an in-house endoscopy dataset demonstrate that DPO-Clin significantly improves the SFT baselines on clinical-aware metrics. Furthermore, it achieves superior performance over existing DPO-based MRG methods, exhibiting robust generalizability across distinct baseline architectures and diverse medical imaging modalities.
CommentsAccepted by ECCV 2026