arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.18988cs.CV

DobicVLM:通过群体相对策略优化使胸部X光报告生成与临床基础的程序化奖励保持一致

DobicVLM: Aligning Chest X-Ray Report Generation with Clinically-Grounded Programmatic Rewards via Group Relative Policy Optimization

  • Dobic Health(多比克健康公司)
  • University of Ibadan(伊巴丹大学)

机构由 AI 辅助整理,请以论文原文为准。

Thanni Adewuyi, Angelica Obayi, Andem Aniekan, Samuel Okoko, Angel Ezendu, Ephraim Usani, Ademide Animasaun, Philip Chibundu, Christian Maurice, Mary Donald Ess… 展开作者

Thanni Adewuyi, Angelica Obayi, Andem Aniekan, Samuel Okoko, Angel Ezendu, Ephraim Usani, Ademide Animasaun, Philip Chibundu, Christian Maurice, Mary Donald Essien, Oluwaseun Odunsi, Oluwasegun Oguntuase, Abiodun Adereni

AI总结:

研究针对自动胸部X光报告生成难题,提出DobicVLM模型,结合监督微调、GRPO及临床奖励,用可解释规则组件执行标准,经训练和评估,该模型多数标准优于Gemini 2.5 Flash,证明GRPO在资源受限设置中的价值。

AI中文摘要:

医学成像在诊断中至关重要,但自动胸部X光报告生成在结构遵循、解剖完整性和语义忠实度方面存在困难。我们引入了DobicVLM,这是一种视觉语言模型,它将在MedGemma - 4B上的监督微调与群体相对策略优化(GRPO)以及基于临床的程序化奖励相结合。我们的方法使用可解释的、基于规则的奖励组件,包括结构验证、解剖检查表、语义相似性和长度约束,以在无需神经奖励模型的情况下执行临床标准。在来自私人临床数据集的1000个去标识化图像 - 报告对(经伦理批准并符合当地法规)上进行训练后,DobicVLM通过对69个保留病例的盲法专家评审进行评估。DobicVLM在大多数标准上优于Gemini 2.5 Flash,与Gemini 2.5 Flash和MedGemma 4B基线相比,在印象准确性(27.2%)和医学术语(86.5%)方面达到最高,在完整性和转诊方面有轻微权衡。这证明了GRPO在资源有限环境中实现透明对齐的价值。

英文摘要:

Medical imaging is a cornerstone of diagnostics, yet automated chest X-ray report generation struggles with structural adherence, anatomical completeness, and semantic faithfulness. We introduce DobicVLM, a vision-language model combining supervised fine-tuning on MedGemma-4B with Group Relative Policy Optimization (GRPO) and clinically-grounded programmatic rewards. Our approach uses interpretable, rule-based reward components; structural verification, anatomical checklist, semantic similarity, and length constraints to enforce clinical standards without neural reward models. Trained on 1,000 de-identified image-report pairs from a private clinical dataset (with ethics approval and compliance to local regulations), DobicVLM is evaluated via blinded expert review on 69 held-out cases. DobicVLM outperforms Gemini 2.5 Flash across the majority of criteria, achieving the highest impression accuracy (27.2%) and medical terminology (86.5%) compared to both Gemini 2.5 Flash and MedGemma 4B baselines, with minor trade-offs in completeness and referrals. This demonstrates GRPO's value for transparent alignment in resource-limited settings. Keywords: Vision-Language Models, Radiology Report Generation, Reinforcement Learning, Medical AI, GRPO

↑