视频推理的帧差分在线策略自蒸馏
Frame Differential On-Policy Self-Distillation for Video Reasoning
浏览论文内容
中文总结 AI 辅助
本文提出帧差分在线策略自蒸馏(FD-OPSD),通过令牌级自蒸馏将密集帧证据迁移至稀疏帧策略,在不改变推理和展开成本的前提下,在六个视频推理基准上超越GRPO等基线。
中文摘要 AI 辅助
强化学习(RL)通过可验证奖励和日益细粒度的视觉或时间信用分配,显著提升了多模态语言模型的推理能力。然而,在视频推理中,当前的RL方法通常使用固定的稀疏帧预算进行训练:增加帧数会使自回归展开成本高昂,而帧数过少则可能遗漏时间局部化事件和细粒度视觉细节。我们提出了帧差分在线策略自蒸馏(FD-OPSD),该方法在RL训练期间将密集帧观测的有用证据迁移到稀疏帧策略中。FD-OPSD比较策略在稀疏和密集视图下对同一采样响应的令牌级偏好,并在没有外部教师或密集自回归展开的情况下蒸馏出所得的帧差分信号。该方法保持了稀疏帧展开,且推理过程不变。在Qwen2.5-VL-7B和Qwen3-VL-4B上,跨六个视频推理基准,FD-OPSD在16、32和64帧评估设置下,相较于最强的对应GRPO、T-GRPO或Video-KTR基线,取得了更高的总体平均性能。这些结果表明,密集视觉证据可以在训练期间通过令牌级自蒸馏选择性迁移,同时保持稀疏帧展开和不变的推理。
英文摘要
Reinforcement learning (RL) has substantially improved the reasoning ability of multimodal language models through verifiable rewards and increasingly fine-grainedvisual or temporal credit assignment. In video reasoning, however, current RL methods typically train with a fixed sparse frame budget: increasing the number of frames makes autoregressive rollouts expensive, while too few frames may miss temporally localized events and fine-grained visual details. We present \textbf{Frame Differential On-Policy Self-Distillation (FD-OPSD)}, which transfers the useful evidence of dense frame observations to a sparse frame policy during RL training. FD-OPSD compares the policy's token level preferences for the same sampled response under sparse and dense views, and distills the resulting frame differential signal without an external teacher or dense autoregressive rollout. The method preserves sparse-frame rollouts and leaves inference unchanged. Across Qwen2.5-VL-7B and Qwen3-VL-4B on six video reasoning benchmarks, FD-OPSD yields higher overall average performance than the strongest corresponding GRPO, T-GRPO, or Video-KTR baselines across the 16, 32, and 64 frame evaluation settings. These results show that dense visual evidence can be transferred selectively during training through token level self-distillation while retaining sparse frame rollouts and unchanged inference.
发表机构
- HKUST(香港科技大学)
- NKU(南开大学)
- SEU(东南大学)
- KAUST(阿卜杜拉国王科技大学)
机构由 AI 辅助整理,请以论文原文为准。