可验证奖励教会视频语言模型关于时间的什么?一项受控多模型研究
What Do Verifiable Rewards Teach Video-Language Models About Time? A Controlled Multi-Model Study
浏览论文内容
中文总结 AI 辅助
本研究通过多模型受控实验发现,可验证奖励训练虽提升视频问答基准准确率,但增益多源于非视觉线索,且不带来时序理解,混合真实视频可防止域外退化。
中文摘要 AI 辅助
基于可验证奖励的强化学习(RLVR)已在语言模型中产生了显著的推理能力提升,而可验证的视频基准使其适用于因果时序视频问答。我们研究RLVR教会视频语言模型关于时间的什么。我们使用组相对策略优化(GRPO)在三种数据方案下微调四个开放模型(Qwen3-VL-8B/4B、Qwen2.5-VL-7B、Gemma-3-12B):已验证(合成的CLEVRER问题,带有精确答案和事件顺序奖励)、未验证(在43,751个真实网络视频上的自监督 pretext 任务)以及1:1混合方案,另加一个已验证+真实分支,在真实视频上增加4,000个可验证问题。每个单元在域内和域外真实视频(NExT-QA时序压力测试集和MVBench子集)上评估,帧按顺序、打乱和缺失三种情况。(1)已验证训练带来较大的域内增益,但随基础能力增强而缩小(较弱模型上+14至+19个百分点;最强模型上+6)。(2)大部分增益是非视觉的:无帧时的准确率提升几乎与有帧时相当。(3)仅已验证训练可严重降低域外准确率,且训练过程中无任何迹象:Qwen3-VL-8B在两个真实视频集上分别损失26.7和25.2个百分点,而混合方案从未显著降低训练模型的性能。添加真实可验证问题消除了该损失(-2.3个百分点,在基础模型噪声范围内),并保持+9.3的域内增益,因此原因是狭窄的仅合成数据,而非验证本身。(4)没有任何方案诱导时序顺序接地:在41项评估中,有序与打乱之间的差距在39项中与零无法区分,在2项中仅边缘显著,尽管存在事件顺序奖励。可验证奖励提高基准准确率,但不带来时序理解。报告无帧对照,并混合真实视频以防域外性能下降。
英文摘要
Reinforcement learning from verifiable rewards (RLVR) has produced large reasoning gains in language models, and verifiable video benchmarks make it applicable to causal-temporal video question answering. We study what RLVR teaches video-language models about time. We fine-tune four open models (Qwen3-VL-8B/4B, Qwen2.5-VL-7B, Gemma-3-12B) with group relative policy optimization under three data recipes: verified (synthetic CLEVRER questions with exact answer and event-order rewards), unverified (self-supervised pretext tasks over 43,751 real web videos), and a 1:1 mixture, plus a verified+real arm that adds 4,000 verifiable questions on real video. Each cell is evaluated in-domain and on out-of-domain real video (a NExT-QA temporal stress set and an MVBench subset), with frames in order, shuffled, and absent. (1) Verified training yields large in-domain gains that shrink as base competence grows (+14 to +19 points on weaker models; +6 on the strongest). (2) Much of the gain is non-visual: accuracy with no frames rises nearly as much as with frames. (3) Verified-only training can severely degrade out-of-domain accuracy with no sign during training: Qwen3-VL-8B loses 26.7 and 25.2 points on the two real-video sets, while the mixture never significantly degrades a model trained on it. Adding real verified questions removes that loss (-2.3 points, within noise of base) and keeps a +9.3 in-domain gain, so the cause is narrow synthetic-only data, not verification. (4) No recipe induces temporal-order grounding: across 41 evaluations the ordered-versus-shuffled gap is indistinguishable from zero in 39 and marginal in two, despite an event-order reward. Verifiable rewards improve benchmark accuracy without temporal understanding. Report no-frame controls, and mix in real video to guard against out-of-domain degradation.
发表机构
- University of Colorado Boulder(科罗拉多大学博尔德分校)
机构由 AI 辅助整理,请以论文原文为准。