发表机构
Zhejiang University; Huazhong University of Science and Technology; Shanghai Jiao Tong University(浙江大学; 华中科技大学; 上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多模态RL中视觉主张缺乏证据支持的问题,提出持久性感知信用门控(PACG),通过衰减异常持久主张的信用,在多个基准上提升准确率并增强主张可撤回性。
AI 中文摘要
可验证奖励强化学习(RLVR)已扩展到大型视觉语言模型(LVLMs),感知感知方法进一步鼓励策略依赖视觉证据。然而,依赖图像并不能保证视觉主张得到图像支持。在RL训练之前,Qwen2.5-VL-7B在四个多模态推理基准上正确回答的响应中,有27.81%包含至少一个图像不支持的直接视觉主张。由于结果级RL将每个响应作为一个整体进行奖励,这些主张继承了正确答案的正向信用。我们引入了一种固定轨迹反事实诊断方法,在干预图像下对同一响应重新评分,以区分证据功能敏感性(EFS,即模型预测变化的强度)与主张持久性(即模型是否继续支持同一主张而非撤回它)。该诊断揭示了敏感性-持久性解耦(SPD):在DAPO和VPPO下,EFS增加且主张整体变得更可撤回,但不受支持的主张变得显著更持久,而GRPO在提高EFS的同时没有这种恶化。因此,我们提出了持久性感知信用门控(PACG),它衰减异常持久的视觉主张的正向信用,并保持所有其他信用不变。它不需要支持/不支持标签,且不增加推理成本。在Qwen2.5-VL-7B上,PACG将三个种子的九个基准平均值从58.1%提高到59.9%(使用DAPO),从59.8%提高到60.9%(使用VPPO),同时使不受支持的主张更可撤回。这些收益扩展到更大的模型、更新的骨干网络,且HallusionBench的准确性也持续提高。这些结果表明,视觉敏感性和主张可撤回性是多模态信用分配的互补维度。
英文摘要
Reinforcement Learning with Verifiable Rewards (RLVR) has been extended to Large Vision-Language Models (LVLMs), and perception-aware methods further encourage policies to rely on visual evidence. Yet relying on the image does not guarantee that visual claims are supported by it. Before RL training, 27.81% of the correctly answered responses of Qwen2.5-VL-7B on four multimodal reasoning benchmarks contain at least one direct visual claim that the image does not support. Since outcome-level RL rewards each response as a whole, these claims inherit the positive credit of the correct answer. We introduce a fixed-rollout counterfactual diagnostic that re-scores the same response under an intervened image to separate Evidence-Function Sensitivity (EFS), how strongly the model's predictions change, from claim persistence, whether the model keeps supporting the same claim rather than retracting it. The diagnostic reveals Sensitivity-Persistence Decoupling (SPD): under DAPO and VPPO, EFS increases and claims become more retractable overall, yet unsupported claims become significantly more persistent, whereas GRPO raises EFS without this deterioration. We therefore propose Persistence-Aware Credit Gating (PACG), which attenuates positive credit for unusually persistent visual claims and leaves all other credit unchanged. It requires no supported/unsupported labels and adds no inference cost. On Qwen2.5-VL-7B, PACG raises the nine-benchmark average over three seeds from 58.1% to 59.9% with DAPO and from 59.8% to 60.9% with VPPO, while making unsupported claims more retractable. The gains extend to a larger model, a newer backbone, and the accuracy of HallusionBench also improves consistently. These results suggest that visual sensitivity and claim retractability are complementary dimensions of multimodal credit assignment.
CommentsMLLM,RL