发表机构
National University of Singapore; DAMO Academy Alibaba Group; Hupan Lab; Zhejiang University; University of California Berkeley; Rochester Institute of Technology; Georgia Institute of Technology; Renmin University of China; Hong Kong University of Science and Technology(新加坡国立大学; 阿里巴巴达摩院; 湖畔实验室; 浙江大学; 加州大学伯克利分校; 罗切斯特理工学院; 佐治亚理工学院; 中国人民大学; 香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多模态RLVR泛化鲁棒性不足的问题,提出含动态三元奖励与一致性正则化器的PIRL方法,压力测试与动态评估中其性能下降幅度均小于GRPO。
AI 中文摘要
带可验证奖励的强化学习(RLVR)可提升多模态大语言模型的准确性,但其增益脆弱:仅对问题进行意译或修改提示模板就会导致性能下降,这对医疗视觉问答(VQA)等高风险场景的可靠部署构成挑战。我们将此归因于标准RL目标的两个问题:其一,二元验证器混淆了格式与内容,奖励信号无法区分错误答案与格式错误;其二,训练分布仅覆盖模型部署时可能遇到的真实提示的一小部分,因此在训练分布上表现良好的策略在测试时遇到未见过的提示时会表现不同。这两种失败都需要一种鲁棒的后训练方法,以帮助策略覆盖更广泛的语义等价提示分布,我们确定了两种有助于实现该目标的措施:在奖励中分离格式与语义,以及对语义等价的扰动提示应用策略不变性。因此,我们提出了提示不变RLVR(PIRL),它由动态三元奖励和基于嵌入空间对抗的一致性正则化器组成。在压力测试中,PIRL在基准上的平均准确率仅下降≤1%,而GRPO下降约3%;在动态评估中,PIRL也实现了最小的性能下降。
英文摘要
Reinforcement Learning with Verifiable Rewards (RLVR) makes Multimodal Large Language Models more accurate, but the gains are brittle: simply paraphrasing a question or changing the prompt template can degrade them, which challenges reliable deployment in high-stakes scenarios like medical VQA. We trace this to two issues of the standard RL objective. First, the binary verifier conflates format with content, so the reward signal cannot tell a wrong answer apart from a misformatted one. Second, the training distribution covers only a thin slice of the real-world prompts that the model might meet at deployment, so policies that perform well on the training distribution can behave differently under unseen prompts during test. Both failures call for a robust post-training method that helps the policy cover a broader distribution of semantically equivalent prompts, and we identify two measures that help achieve this objective: separating format from semantics in the reward, and applying policy invariance across perturbed prompts with equivalent semantics. We therefore propose Prompt-Invariant RLVR (PIRL), consisting of a dynamic trinary reward and a consistency regularizer based on an embedding-space adversary. Under stress testing, PIRL's average accuracy on benchmarks drops by only $\le 1\%$, where GRPO drops ~3%. On dynamic evaluation, PIRL also achieves the smallest performance drop.
Comments32 pages, 5 figures