对齐思考与答案:概率奖励以驯服思考漂移
Aligning Thoughts with Answers: Probability Rewards to Tame Thinking Drift
浏览论文内容
中文总结 AI 辅助
本文针对视觉语言模型中的思考漂移问题,提出Rita强化学习范式,通过基于答案条件概率的思考与一致性奖励及难度感知数据过滤,在EgoIntention和RefEgo-Int基准上显著提升性能。
中文摘要 AI 辅助
本文研究了视觉语言模型中的“思考-答案一致性”问题。我们聚焦于视觉意图定位任务,该任务要求模型根据人类意图查询推断目标对象并预测边界框。我们发现,先前基于IoU的强化学习(RL)框架存在“思考漂移”问题,即模型虽然输出了正确的边界框,但其推理过程却指向了不同的目标对象。为此,我们提出了Rita(ReInforcing Thinking--Answer consistency),一种新颖的RL范式,以驯服这种漂移。具体而言,Rita引入了两种无需推理标签的RL奖励,它们基于参考答案的条件概率构建:思考奖励和一致性奖励。此外,Rita还采用了一种难度感知的数据过滤策略,利用 rollout 错误率和奖励方差选择信息量丰富的易到中等难度样本用于RL训练。在EgoIntention和新的RefEgo-Int基准上的大量实验表明,Rita在性能上持续优于监督微调方法和普通RL微调框架。
英文摘要
This paper studies \textbf{thinking--answer consistency} in vision-language models. We focus on Visual Intention Grounding, where a model infers a target object based on a human intention query and predicts a bounding box. We reveal that previous IoU-based reinforcement learning (RL) frameworks suffer from ``thinking drift'', where the model produces a correct bounding box, despite having an incorrect reasoning process pointing to a different target object. Thus, we propose \textbf{Rita} (\textit{ReInforcing Thinking--Answer consistency}) as a novel RL paradigm to tame the drift. Specifically, Rita introduces two reasoning-label-free RL rewards, constructed from the conditional probability of reference answers: a \textbf{thinking reward} and a \textbf{consistency reward}. It also adopts a difficulty-aware \textbf{data filtering} strategy that selects informative easy-to-medium samples for RL using rollout error rate and reward variance. Extensive experiments on EgoIntention and the new RefEgo-Int benchmarks show that Rita performs consistently superior to the supervised finetuning approaches and vanilla RL-finetuned frameworks.
发表机构
- National University of Singapore(新加坡国立大学)
- A*STAR Institute of Advanced Intelligence and Computing, Singapore(新加坡科技研究局先进智能与计算研究所)
- South China University of Technology(华南理工大学)
- University of Science and Technology of China(中国科学技术大学)
- Google DeepMind(谷歌DeepMind)
机构由 AI 辅助整理,请以论文原文为准。