arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.13931cs.CV

SIVA-RL:用于多模态强化学习的灵敏度不变视觉对齐

SIVA-RL: Sensitivity-Invariance Visual Alignment for Multimodal Reinforcement Learning

Cheng Tang, Junzhi Ning, Min Cen, Wei Li, Xinyi Zeng, Pinxian Zeng, Rongbin Li, Qiming Zhu, Yuqiang Li, Junjun He, Yirong Chen, Ming Hu

首次发表
浏览论文内容

中文总结 AI 辅助

研究多模态强化学习中视觉语言模型预测与视觉证据结合问题,提出SIVA-RL框架,通过特定方法构建局部干预并以奖励下降为权重驱动对齐,在多基准测试中相比基线改进了模型。

中文摘要 AI 辅助

具有可验证奖励的强化学习(RLVR)推动多模态推理,但答案级别的正确性并不能保证视觉语言模型的预测基于视觉证据。现有的视觉干预方法对比原始图像和修改后图像上的策略行为,但按干预类型而非观察到的效果分配监督。我们提出了SIVA-RL,一种灵敏度不变视觉对齐框架,用逐样本、基于结果的监督取代基于算子的正则化。SIVA-RL通过令牌对齐、距离受限的图像内PatchSwap构建局部干预。然后,一个冻结的审计策略对每个干净-干预对进行评分,观察到的奖励下降成为软路由权重。大下降对驱动灵敏度对齐,小下降对驱动干净锚定的不变性对齐,模糊对权重降低。该设计将干预构建与监督分配解耦,与GRPO和DAPO主干兼容。在九个多模态推理基准测试中,SIVA-RL在每种设置下都比匹配的RL基线改进了3B和7B模型。在基于视觉的推理上提高了8.79个百分点,在所有四种基于GRPO和DAPO的配置中总体相对提高了14.9%。

英文摘要

Reinforcement learning with verifiable rewards (RLVR) drives multimodal reasoning, but answer-level correctness does not guarantee that a vision-language model grounds its predictions in visual evidence. Existing visual-intervention methods contrast policy behavior on original and modified images, yet assign supervision by the type of intervention rather than its observed effect. This assumption fails: identical operators produce heterogeneous outcomes across samples. We propose SIVA-RL, a Sensitivity-Invariance Visual Alignment framework that replaces operator-conditioned regularization with sample-wise, outcome-conditioned supervision. SIVA-RL constructs localized interventions through token-aligned, distance-constrained within-image PatchSwap. A frozen audit policy then scores each clean-intervention pair, and the observed reward drop becomes soft routing weights. Large-drop pairs drive sensitivity alignment, low-drop pairs drive clean-anchored invariance alignment, and ambiguous pairs are down-weighted. This design decouples intervention construction from supervision assignment and is compatible with both GRPO and DAPO backbones. Across nine multimodal reasoning benchmarks spanning mathematical, logical, and vision-dependent tasks, SIVA-RL improves 3B and 7B models over matched RL baselines in every setting. It yields an 8.79 percentage-point gain on vision-dependent reasoning and up to 14.9% relative overall improvement across all four GRPO- and DAPO-based configurations.

发表机构

  • Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
  • Shanghai Jiao Tong University(上海交通大学)
  • Sichuan University(四川大学)
  • The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
  • University of Macau(澳门大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑