arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39183cs.CV

对齐思考与答案:概率奖励以驯服思考漂移

Aligning Thoughts with Answers: Probability Rewards to Tame Thinking Drift

Pengzhan Sun, Shiu-hong Kao, Shijie Li, Yongyi Su, Junbin Xiao, Arjun Reddy Akula, Angela Yao

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对视觉语言模型中的思考漂移问题,提出Rita强化学习范式,通过基于答案条件概率的思考与一致性奖励及难度感知数据过滤,在EgoIntention和RefEgo-Int基准上显著提升性能。

中文摘要 AI 辅助

本文研究了视觉语言模型中的“思考-答案一致性”问题。我们聚焦于视觉意图定位任务,该任务要求模型根据人类意图查询推断目标对象并预测边界框。我们发现,先前基于IoU的强化学习(RL)框架存在“思考漂移”问题,即模型虽然输出了正确的边界框,但其推理过程却指向了不同的目标对象。为此,我们提出了Rita(ReInforcing Thinking--Answer consistency),一种新颖的RL范式,以驯服这种漂移。具体而言,Rita引入了两种无需推理标签的RL奖励,它们基于参考答案的条件概率构建:思考奖励和一致性奖励。此外,Rita还采用了一种难度感知的数据过滤策略,利用 rollout 错误率和奖励方差选择信息量丰富的易到中等难度样本用于RL训练。在EgoIntention和新的RefEgo-Int基准上的大量实验表明,Rita在性能上持续优于监督微调方法和普通RL微调框架。

英文摘要

This paper studies \textbf{thinking--answer consistency} in vision-language models. We focus on Visual Intention Grounding, where a model infers a target object based on a human intention query and predicts a bounding box. We reveal that previous IoU-based reinforcement learning (RL) frameworks suffer from ``thinking drift'', where the model produces a correct bounding box, despite having an incorrect reasoning process pointing to a different target object. Thus, we propose \textbf{Rita} (\textit{ReInforcing Thinking--Answer consistency}) as a novel RL paradigm to tame the drift. Specifically, Rita introduces two reasoning-label-free RL rewards, constructed from the conditional probability of reference answers: a \textbf{thinking reward} and a \textbf{consistency reward}. It also adopts a difficulty-aware \textbf{data filtering} strategy that selects informative easy-to-medium samples for RL using rollout error rate and reward variance. Extensive experiments on EgoIntention and the new RefEgo-Int benchmarks show that Rita performs consistently superior to the supervised finetuning approaches and vanilla RL-finetuned frameworks.

发表机构

  • National University of Singapore(新加坡国立大学)
  • A*STAR Institute of Advanced Intelligence and Computing, Singapore(新加坡科技研究局先进智能与计算研究所)
  • South China University of Technology(华南理工大学)
  • University of Science and Technology of China(中国科学技术大学)
  • Google DeepMind(谷歌DeepMind)

机构由 AI 辅助整理,请以论文原文为准。

↑