Video-HopChain:面向视频推理模型的多跳问题与置信门控探索
Video-HopChain: Multi-Hop Questions and Confidence-Gated Exploration for Video Reasoning Models
浏览论文内容
中文总结 AI 辅助
针对视频推理模型,构建多跳问答数据集Video-HopChain,并引入置信门控探索(CGE)以克服GRPO在难题和易题上的梯度消失,显著提升多基准平均分。
中文摘要 AI 辅助
HopChain已在静态图像上证明,多跳数据合成能够提升视觉-语言推理能力,因为长链式思维推理会暴露跨步骤累积的错误,而用于可验证奖励强化学习(RLVR)的大多数数据很少要求视觉证据链,因此这些弱点很可能未被暴露。我们在视频中观察到同样的问题,而该框架在视频领域尚未被探索。为此,我们构建了Video-HopChain数据集,包含13,378个视频上的22,550个多跳视频问题,以及一个留出的1,000个问题的基准测试集。每个问题将关于同一视频中多个时刻的三至六个是非题串联起来,每个问题根据其答案产生两个整数之一。最终答案是这些整数之和,因此对该和的精确匹配提供了RLVR所需的可验证奖励。我们首先使用GRPO在标准视频数据集上训练Qwen3-VL-8B,然后在Video-HopChain上进行第二阶段训练,将八个视频理解与推理基准的平均分从55.4提升至57.9,并且每个基准都有所改善。然而,在此类数据集上训练暴露了GRPO的一个已知局限:其学习信号来自组内奖励方差,因此所有rollout均错误的难题和所有rollout均正确的易题都会使组内无梯度。为了在相同计算预算下恢复这些组,我们引入了置信门控探索(CGE)。每个问题使用8次rollout时,CGE按常规采样前4次。如果这4次要么全部正确要么全部错误,则使用策略在推理跨度内被掩码的最自信token来采样后4次,并从损失中移除被掩码的位置,同时所有8次rollout都进入优势计算。使用CGE后,平均分进一步升至59.3。我们发布了数据集、检查点以及数据生成和训练代码。
英文摘要
HopChain has shown on still images that multi-hop data synthesis improves vision-language reasoning, because long chain-of-thought reasoning exposes errors that compound across steps, while most data used for reinforcement learning with verifiable rewards (RLVR) rarely demands a chain of visual evidence, so these weaknesses are likely to stay unexposed. We observe the same problem in video, where this framework has not yet been explored. We therefore build Video-HopChain, a dataset of 22,550 multi-hop video questions over 13,378 videos, together with a held-out benchmark of 1,000 questions. Each question chains three to six yes/no questions about moments in one video, and each yields one of two integers depending on its answer. The final answer is the sum of these integers, so an exact match on that sum gives the verifiable reward that RLVR needs. We first train Qwen3-VL-8B with GRPO on a standard video dataset, and a second stage on Video-HopChain then raises the mean over eight video understanding and reasoning benchmarks from 55.4 to 57.9 and improves every one of them. Training on such a dataset, however, exposes a known limitation of GRPO: its learning signal comes from the reward variance within a group, so hard questions whose rollouts are all incorrect and easy questions whose rollouts are all correct both leave the group with no gradient. To recover these groups at the same compute budget, we introduce Confidence-Gated Exploration (CGE). With 8 rollouts per question, CGE samples the first 4 as usual. If these 4 are either all correct or all incorrect, it samples the last 4 with the policy's most confident token masked inside the reasoning span, and removes the masked positions from the loss while all 8 rollouts enter the advantage. With CGE, the mean rises further to 59.3. We release the dataset, the checkpoint, and the data generation and training code.
发表机构
- Nanyang Technological University (NTU)(南洋理工大学)
- Johns Hopkins University(约翰斯·霍普金斯大学)
机构由 AI 辅助整理,请以论文原文为准。