arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

残差视觉信用优化:用于多模态强化学习的守恒证据路由

Residual Visual Credit Optimization: Conserved Evidence Routing for Multimodal Reinforcement Learning

Lin Qiu, Yao Liu, Diyi Hu, Hanqing Zeng, Onur Gungor, Chujie Chen, Jiayi Liu, Jianyu Wang, XueLin Zheng

arXiv 2610.04918首次发表:更新:

发表机构

Meta Superintelligence Labs; University of Southern California; Meta Recommendation System; Meta(Meta超级智能实验室; 南加州大学; Meta推荐系统; Meta)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多模态强化学习中结果奖励无法分配至各决策的问题,提出残差视觉信用优化(RVCO),通过守恒证据路由将信用按感知依赖分配,在四个模型家族和七个基准上提升准确率并保持训练稳定性。

AI 中文摘要

基于可验证奖励的强化学习能够扩展多模态推理能力,但结果奖励仅表明整个轨迹的价值,并未说明该价值应如何在产生它的各个决策之间分配。我们提出了残差视觉信用优化(RVCO),将令牌信用视为一个守恒路由问题。受控的视觉干预产生每个令牌的证据响应;轨迹内稳健坐标消除偶然的尺度差异;一个有预算的熵路由器根据感知依赖将固定量的序列效用进行分配。残差支持路径保证每个有效位置获得正信用,解析校正精确恢复规定的信用质量。由此产生的信用场具有选择性、有界性、全支持性,并且对响应局部分数平移保持不变,同时将硬令牌选择恢复为极限情况。在四个模型家族和七个推理基准上,RVCO在强RLVR基线上提高了准确率,同时保持了后期优化稳定性、损坏鲁棒性和具有竞争力的训练成本。奖励、轨迹采样和组相对优势估计器保持不变;仅令牌级信用的几何结构不同。

英文摘要

Reinforcement learning with verifiable rewards scales multimodal reasoning, but an outcome reward says how much a trajectory is worth, not how that value should be spread over the decisions that produced it. We introduce Residual Visual Credit Optimization (RVCO), which treats token credit as a conserved routing problem. A controlled visual intervention yields a per-token evidence response; robust within-trajectory coordinates remove incidental scale; and a budgeted entropic router distributes a fixed amount of sequence utility according to perceptual dependence. A residual support path guarantees positive credit at every valid position, and an analytic correction restores the prescribed credit mass exactly. The resulting field is selective, bounded, full-support, and invariant to response-local score shifts, and recovers hard token selection as a limiting case. Across four model families and seven reasoning benchmarks, RVCO improves accuracy over strong RLVR baselines while maintaining late-stage optimization stability, corruption robustness, and competitive training cost. Rewards, rollouts, and the group-relative advantage estimator are unchanged; only the geometry of token-level credit differs.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑