arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.10147cs.CV

REVA-PO:用于胸部X光报告生成的稳定强化学习

REVA-PO: Stabilizing Reinforcement Learning for Chest X-ray Report Generation

发表机构英属哥伦比亚大学
查看机构详情
  • University of British Columbia(英属哥伦比亚大学)

机构由 AI 辅助整理,请以论文原文为准。

Li Guo, Anas M. Tahir, Z. Jane Wang

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对胸部X光报告生成中强化学习训练不稳定等问题,提出REVA-PO框架,通过响应加权正则化和验证锚定策略重置稳定训练,采用三阶段管道,实验证明该框架在语言质量和临床准确性上达新基准。

中文摘要 AI 辅助

自动胸部X光报告生成最近受益于强化学习(RL)和大语言模型。然而,由于固定的Kullback-Leibler(KL)正则化和随着时间积累KL压力的静态参考策略,RL训练常常存在不稳定性或探索受限的问题。我们提出了响应加权和验证锚定策略优化(REVA-PO),这是一个通过响应加权正则化(RER)和验证锚定策略重置(VAPR)来稳定长期训练的RL框架。RER根据优势和参考策略熵动态调整每个响应的KL权重,VAPR定期将参考和当前策略同步到最佳验证检查点。我们采用了由热身训练、分类器引导的监督微调以及RL组成的三阶段管道。在MIMIC-CXR和IU-Xray上的广泛评估表明,REVA-PO在语言质量和临床准确性方面都设定了新的最先进基准。

英文摘要

Automated chest X-ray report generation has recently benefited from reinforcement learning (RL) and large language models. However, RL training often suffers from instability or limited exploration due to fixed Kullback-Leibler (KL) regularization and a static reference policy that accumulates KL pressure over time. We propose Response-Weighted and Validation-Anchored Policy Optimization (REVA-PO), a RL framework that stabilizes long-term training via Response-Weighted Regularization (RER) and Validation-Anchored Policy Reset (VAPR). RER dynamically adjusts per-response KL weights based on advantage and reference-policy entropy, relaxing constraints for high-quality responses while tightening them for low-quality ones. Complementarily, VAPR periodically synchronizes the reference and current policies to the best validation checkpoint, resetting accumulated regularization pressure to expand the viable exploration space. To ensure a robust starting point, we employ a three-stage pipeline consisting of warm-up training, classifier-guided supervised fine-tuning, and RL. Extensive evaluations on MIMIC-CXR and IU-Xray demonstrate that REVA-PO sets new state-of-the-art benchmarks in both linguistic quality and clinical accuracy. Notably, BLEU-4 improves by 5.1% on MIMIC-CXR and 3.6% on IU-Xray, while CheXpert F1 and RadGraph F1 scores increase by 4.5% and 12.8%, respectively, over prior leading methods. The code is publicly available at https://github.com/LiGuo12/REVA_PO/.

补充信息

↑