REVA-PO:用于胸部X光报告生成的稳定强化学习
REVA-PO: Stabilizing Reinforcement Learning for Chest X-ray Report Generation
查看机构详情
- University of British Columbia(英属哥伦比亚大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
研究针对胸部X光报告生成中强化学习训练不稳定等问题,提出REVA-PO框架,通过响应加权正则化和验证锚定策略重置稳定训练,采用三阶段管道,实验证明该框架在语言质量和临床准确性上达新基准。
中文摘要 AI 辅助
自动胸部X光报告生成最近受益于强化学习(RL)和大语言模型。然而,由于固定的Kullback-Leibler(KL)正则化和随着时间积累KL压力的静态参考策略,RL训练常常存在不稳定性或探索受限的问题。我们提出了响应加权和验证锚定策略优化(REVA-PO),这是一个通过响应加权正则化(RER)和验证锚定策略重置(VAPR)来稳定长期训练的RL框架。RER根据优势和参考策略熵动态调整每个响应的KL权重,VAPR定期将参考和当前策略同步到最佳验证检查点。我们采用了由热身训练、分类器引导的监督微调以及RL组成的三阶段管道。在MIMIC-CXR和IU-Xray上的广泛评估表明,REVA-PO在语言质量和临床准确性方面都设定了新的最先进基准。
英文摘要
Automated chest X-ray report generation has recently benefited from reinforcement learning (RL) and large language models. However, RL training often suffers from instability or limited exploration due to fixed Kullback-Leibler (KL) regularization and a static reference policy that accumulates KL pressure over time. We propose Response-Weighted and Validation-Anchored Policy Optimization (REVA-PO), a RL framework that stabilizes long-term training via Response-Weighted Regularization (RER) and Validation-Anchored Policy Reset (VAPR). RER dynamically adjusts per-response KL weights based on advantage and reference-policy entropy, relaxing constraints for high-quality responses while tightening them for low-quality ones. Complementarily, VAPR periodically synchronizes the reference and current policies to the best validation checkpoint, resetting accumulated regularization pressure to expand the viable exploration space. To ensure a robust starting point, we employ a three-stage pipeline consisting of warm-up training, classifier-guided supervised fine-tuning, and RL. Extensive evaluations on MIMIC-CXR and IU-Xray demonstrate that REVA-PO sets new state-of-the-art benchmarks in both linguistic quality and clinical accuracy. Notably, BLEU-4 improves by 5.1% on MIMIC-CXR and 3.6% on IU-Xray, while CheXpert F1 and RadGraph F1 scores increase by 4.5% and 12.8%, respectively, over prior leading methods. The code is publicly available at https://github.com/LiGuo12/REVA_PO/.