SIVA-RL:用于多模态强化学习的灵敏度不变视觉对齐
SIVA-RL: Sensitivity-Invariance Visual Alignment for Multimodal Reinforcement Learning
浏览论文内容
中文总结 AI 辅助
研究多模态强化学习中视觉语言模型预测与视觉证据结合问题,提出SIVA-RL框架,通过特定方法构建局部干预并以奖励下降为权重驱动对齐,在多基准测试中相比基线改进了模型。
中文摘要 AI 辅助
具有可验证奖励的强化学习(RLVR)推动多模态推理,但答案级别的正确性并不能保证视觉语言模型的预测基于视觉证据。现有的视觉干预方法对比原始图像和修改后图像上的策略行为,但按干预类型而非观察到的效果分配监督。我们提出了SIVA-RL,一种灵敏度不变视觉对齐框架,用逐样本、基于结果的监督取代基于算子的正则化。SIVA-RL通过令牌对齐、距离受限的图像内PatchSwap构建局部干预。然后,一个冻结的审计策略对每个干净-干预对进行评分,观察到的奖励下降成为软路由权重。大下降对驱动灵敏度对齐,小下降对驱动干净锚定的不变性对齐,模糊对权重降低。该设计将干预构建与监督分配解耦,与GRPO和DAPO主干兼容。在九个多模态推理基准测试中,SIVA-RL在每种设置下都比匹配的RL基线改进了3B和7B模型。在基于视觉的推理上提高了8.79个百分点,在所有四种基于GRPO和DAPO的配置中总体相对提高了14.9%。
英文摘要
Reinforcement learning with verifiable rewards (RLVR) drives multimodal reasoning, but answer-level correctness does not guarantee that a vision-language model grounds its predictions in visual evidence. Existing visual-intervention methods contrast policy behavior on original and modified images, yet assign supervision by the type of intervention rather than its observed effect. This assumption fails: identical operators produce heterogeneous outcomes across samples. We propose SIVA-RL, a Sensitivity-Invariance Visual Alignment framework that replaces operator-conditioned regularization with sample-wise, outcome-conditioned supervision. SIVA-RL constructs localized interventions through token-aligned, distance-constrained within-image PatchSwap. A frozen audit policy then scores each clean-intervention pair, and the observed reward drop becomes soft routing weights. Large-drop pairs drive sensitivity alignment, low-drop pairs drive clean-anchored invariance alignment, and ambiguous pairs are down-weighted. This design decouples intervention construction from supervision assignment and is compatible with both GRPO and DAPO backbones. Across nine multimodal reasoning benchmarks spanning mathematical, logical, and vision-dependent tasks, SIVA-RL improves 3B and 7B models over matched RL baselines in every setting. It yields an 8.79 percentage-point gain on vision-dependent reasoning and up to 14.9% relative overall improvement across all four GRPO- and DAPO-based configurations.
发表机构
- Shanghai Artificial Intelligence Laboratory(上海人工智能实验室)
- Shanghai Jiao Tong University(上海交通大学)
- Sichuan University(四川大学)
- The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
- University of Macau(澳门大学)
机构由 AI 辅助整理,请以论文原文为准。