arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

V-Rubrics:基于评分规则的强化学习实现视觉忠实性

V-Rubrics: Visual Faithfulness via Rubric-Based Reinforcement Learning

Shulin Tian, Minglun Li, Yuhao Dong, Hao Ding, Jiarui Yao, Haiwen Diao, Jingkang Yang, Hongyuan Zhu, Ziwei Liu

arXiv 2608.25580首次发表:更新:

发表机构

S-Lab, Nanyang Technological University; UIUC; A*STAR(新加坡南洋理工大学S-Lab; 伊利诺伊大学厄巴纳-香槟分校; 新加坡科技研究局)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对视觉-语言模型回答缺乏视觉依据的问题,提出基于评分规则的强化学习方法V-Rubrics,通过分解响应为原子命题并从三方面评分,训练的模型在相关推理基准上性能优于基线。

AI 中文摘要

视觉-语言模型能够生成流畅的回答,但这些回答可能缺乏对视觉证据的充分依据:单个未被支撑的物体、图表数值或中间推理都可能破坏原本看似合理的响应。我们认为这是多模态后训练中的信用分配失败问题。标量结果奖励仅指示回答是否可接受,但无法确定哪些视觉事实有依据、哪些推理步骤有效,或哪些指令约束未被满足。我们引入了基于视觉评分规则的强化学习(Visual Rubrics-Based Reinforcement Learning),该方法将参考响应分解为原子命题,并沿视觉忠实性(Visual Faithfulness,VF)、推理一致性(Reasoning Consistency,RC)和指令遵循(Instruction Following,IF)对生成的回答进行评分。生成的评分规则项提供结构化的部分信用,并在存在支撑证据区间时定位评分规则信用。我们首先通过在公开的OpenMMReasoner-SFT-874K语料库上微调Qwen3-VL-8B-Instruct获取SFT检查点,适配OpenMMReasoner的冷启动数据方案。我们从17个视觉依据来源构建了包含50248个示例的训练集V-Rubrics 50K,具体步骤为:应用基于规则的过滤器,再通过拒绝采样分数推导示例难度,最后使用Gemini-3-Pro在相同的结构化提示和协议下为每个示例标注。我们基于相同的SFT检查点,使用分量式、前缀定位的评分规则信用训练模型。实验表明,我们的基于评分规则的GRPO(Group Relative Policy Optimization)相较于共享SFT基线和仅回答的GRPO均有提升,在面向知识和视觉依据的推理基准上获得了最大的增益。结果表明,评分规则是视觉后训练的一种有用的奖励抽象。

英文摘要

Vision-language models can produce fluent answers that are insufficiently grounded in the visual evidence: a single unsupported object, chart value, or intermediate inference can undermine an otherwise plausible response. We argue that this is a credit-assignment failure in multimodal post-training. Scalar outcome rewards indicate whether an answer is acceptable, but do not identify which visual facts are grounded, which reasoning steps are valid, or which instruction constraints are missed. We introduce Visual Rubrics-Based Reinforcement Learning, which decomposes reference responses into atomic propositions and scores generated answers along Visual Faithfulness (VF), Reasoning Consistency (RC), and Instruction Following (IF). The resulting rubric items provide structured partial credit and localize rubric credit when supporting evidence spans are available. We first obtain an SFT checkpoint by fine-tuning Qwen3-VL-8B-Instruct on the public OpenMMReasoner-SFT-874K corpus, adapting OpenMMReasoner's cold-start data recipe. We construct V-Rubrics 50K, a 50,248-example training set from 17 visually grounded sources, by applying rule-based filters before deriving example difficulty from rejection-sampling scores and then annotating every example with Gemini-3-Pro under the same structured prompt and protocol. We train our model based on the same SFT checkpoint using component-wise, prefix-localized rubric credit. Experiments show that our rubricbased GRPO improves over both the shared SFT baseline and answer-only GRPO, with the largest gains on knowledge-oriented and visually grounded reasoning benchmarks. The results show rubrics as a useful reward abstraction for visual post-training.

CommentsProj page: https://shulin16.github.io/v-rubrics/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑