DobicVLM:通过群体相对策略优化使胸部X光报告生成与临床基础的程序化奖励保持一致
DobicVLM: Aligning Chest X-Ray Report Generation with Clinically-Grounded Programmatic Rewards via Group Relative Policy Optimization
- Dobic Health(多比克健康公司)
- University of Ibadan(伊巴丹大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究针对自动胸部X光报告生成难题,提出DobicVLM模型,结合监督微调、GRPO及临床奖励,用可解释规则组件执行标准,经训练和评估,该模型多数标准优于Gemini 2.5 Flash,证明GRPO在资源受限设置中的价值。
AI中文摘要:
医学成像在诊断中至关重要,但自动胸部X光报告生成在结构遵循、解剖完整性和语义忠实度方面存在困难。我们引入了DobicVLM,这是一种视觉语言模型,它将在MedGemma - 4B上的监督微调与群体相对策略优化(GRPO)以及基于临床的程序化奖励相结合。我们的方法使用可解释的、基于规则的奖励组件,包括结构验证、解剖检查表、语义相似性和长度约束,以在无需神经奖励模型的情况下执行临床标准。在来自私人临床数据集的1000个去标识化图像 - 报告对(经伦理批准并符合当地法规)上进行训练后,DobicVLM通过对69个保留病例的盲法专家评审进行评估。DobicVLM在大多数标准上优于Gemini 2.5 Flash,与Gemini 2.5 Flash和MedGemma 4B基线相比,在印象准确性(27.2%)和医学术语(86.5%)方面达到最高,在完整性和转诊方面有轻微权衡。这证明了GRPO在资源有限环境中实现透明对齐的价值。
英文摘要:
Medical imaging is a cornerstone of diagnostics, yet automated chest X-ray report generation struggles with structural adherence, anatomical completeness, and semantic faithfulness. We introduce DobicVLM, a vision-language model combining supervised fine-tuning on MedGemma-4B with Group Relative Policy Optimization (GRPO) and clinically-grounded programmatic rewards. Our approach uses interpretable, rule-based reward components; structural verification, anatomical checklist, semantic similarity, and length constraints to enforce clinical standards without neural reward models. Trained on 1,000 de-identified image-report pairs from a private clinical dataset (with ethics approval and compliance to local regulations), DobicVLM is evaluated via blinded expert review on 69 held-out cases. DobicVLM outperforms Gemini 2.5 Flash across the majority of criteria, achieving the highest impression accuracy (27.2%) and medical terminology (86.5%) compared to both Gemini 2.5 Flash and MedGemma 4B baselines, with minor trade-offs in completeness and referrals. This demonstrates GRPO's value for transparent alignment in resource-limited settings. Keywords: Vision-Language Models, Radiology Report Generation, Reinforcement Learning, Medical AI, GRPO