arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Token级视频强化学习

Token-Level Video Reinforcement Learning

Yifan Wang, Gordon Guocheng Qian, Yanyu Li, Anil Kag, Yun Fu

arXiv 2610.01973首次发表:更新:

发表机构

Northeastern University(东北大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对视频生成中标量奖励无法定位错误的问题,提出Token级视频强化学习框架,利用冻结视觉语言模型的梯度生成token级信用,在GRPO中重加权去噪概率,在VBench-2.0上提升3.60分。

AI 中文摘要

视频生成的强化学习(RL)通常为整个采样视频分配一个标量奖励。然而,视频并非均匀地存在缺陷:某些视觉token可能已经满足提示要求,而其他token则需要修正。标量奖励无法定位错误,导致优化会扰动满意的token,同时对实际需要改变的token关注不足。我们引入了Token级视频强化学习(TVRL),这是一个从被优化的奖励中推导出token级信用的框架。我们的关键洞察是,冻结的视觉-语言模型的答案似然提供了两种信号:其输出贡献于视频级奖励,而其视频输入梯度的幅度揭示了哪些生成的视频token对该分数影响最大。我们在组相对策略优化中实例化TVRL,通过将提示衍生的问题奖励平均为一个组相对优势,并使用分离的、问题条件化的token信用图来在裁剪的策略比率内重新加权密集的去噪转换对数概率。在VBench-2.0上,TVRL取得了57.69的总体分数,比基础模型高出3.60分。TVRL还在三个SDE采样器(SAGE、Flow和Dance)上比匹配的GRPO基线提高了2.68-3.15分,并在四个奖励模型(VideoAlign、VideoScore2、UnifiedReward2和Qwen3.5-9B)上提高了1.33-3.15分。

英文摘要

Reinforcement learning (RL) for video generation usually assigns one scalar reward to an entire sampled video. Yet a video is not uniformly flawed: some visual tokens may already satisfy the prompt, whereas others require correction. A scalar reward cannot localize errors, causing optimization to perturb satisfactory tokens while under-targeting the tokens that actually need to change. We introduce Token-Level Video Reinforcement Learning, TVRL, a framework that derives token-level credit from the reward being optimized. Our key insight is that the answer likelihood of a frozen vision-language model provides both signals: its outputs contribute to the video-level reward, while magnitudes of its video-input gradients reveal which generated video tokens most affect that score. We instantiate TVRL in Group Relative Policy Optimization by averaging prompt-derived question rewards into one group-relative advantage and using detached, question-conditioned token-credit maps to reweight dense denoising-transition log-probabilities inside the clipped policy ratio. On VBench-2.0, TVRL achieves an Overall score of 57.69, outperforming the base model by 3.60 points. TVRL also improves matched GRPO baselines across three SDE samplers (SAGE, Flow, and Dance) by 2.68--3.15 points and across four reward models (VideoAlign, VideoScore2, UnifiedReward2, and Qwen3.5-9B) by 1.33--3.15 points.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑