arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

利用强化学习增强胸部X光报告生成中的视觉推理能力

Enhancing Visual Reasoning in Chest X-Ray Report Generation Using Reinforcement Learning

Denis Musinguzi, Andrew Katumba, Prasenjit Mitra

arXiv 2609.31911首次发表:更新:

AI 中文总结

本研究提出一种基于强化学习的框架,通过验证解剖区域、边界框和区域级文本描述来增强胸部X光报告生成的视觉推理,减少幻觉并提升临床可审计性。

AI 中文摘要

随着现代视觉语言模型的兴起和大规模医学数据集的日益丰富,医学报告生成取得了显著进展。然而,幻觉现象仍是一个主要挑战,这很大程度上归因于监督微调(SFT)的局限性,它优先考虑与参考报告的词汇相似性,而非临床正确性。虽然强化学习在数学和代码生成等具有可验证奖励的领域表现出色,但其在开放式医学任务中的应用仍然有限。现有工作侧重于评估最终答案,忽视了模型的推理过程,尽管有证据表明有缺陷的推理会降低整体性能。在本研究中,我们提出一个框架,通过整合解剖区域、边界框和区域级文本描述来验证模型的推理过程。我们设计了空间和事实奖励机制,以确保模型的推理既具有视觉基础又具有事实准确性。以Qwen3-VL-8B-Instruct为基础模型,我们通过监督微调使其适应医学领域,通过冷启动SFT阶段引入推理能力,并使用强化学习进行优化。我们发现,强化学习带来的性能提升超出了仅通过SFT所能达到的水平,并且联合验证推理步骤和最终输出比单独验证任一者能带来更大的改进。我们进一步识别了强化学习阶段中多种奖励黑客模式。最后,模型的结构化思考轨迹增强了可解释性,使其输出更易于临床使用的审计。

英文摘要

Medical report generation has made significant progress with the rise of modern vision-language models and the growing availability of large-scale medical datasets. However, hallucinations remain a major challenge, largely due to the limitations of supervised fine-tuning (SFT), which prioritizes lexical similarity to reference reports rather than clinical correctness. While reinforcement learning has shown strong performance in domains with verifiable rewards such as mathematics and code generation, its application to open-ended medical tasks remains limited. Existing work focuses on evaluating final answers, overlooking the model's reasoning, despite evidence that flawed reasoning can degrade overall performance. In this study, we propose a framework that verifies the model's reasoning process by integrating anatomical regions, bounding boxes, and region-level textual descriptions. We design spatial and factual reward mechanisms to ensure that the model's reasoning is both visually grounded and factually accurate. Starting from Qwen3-VL-8B-Instruct as our base model, we adapt it to the medical domain using supervised fine-tuning, introduce reasoning capability through a cold-start SFT stage, and refine it with reinforcement learning. We find that RL provides performance gains beyond those achievable through SFT alone, and that jointly verifying both reasoning steps and final outputs yields larger improvements than verifying either in isolation. We further identify multiple modes of reward hacking in the RL stage. Finally, the model's structured think traces enhance interpretability, making its outputs easier to audit for clinical use.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑