AI 中文总结
本研究质疑视频-语言模型依赖更多时间观测的假设,提出SAVER密集到稀疏后训练框架,仅用1250个时间grounding示例训练,可在少帧下提升稀疏视频推理性能,迁移至视频问答等任务。
AI 中文摘要
视频-语言模型通常假设更多的时间观测会带来更可靠的推理。我们质疑这一假设,并认为关键挑战不仅是高效处理更多视频帧,而是在有限的时间证据下学习可靠推理。我们提出SAVER,一种密集到稀疏的后训练框架,该框架在训练时使用密集视频视图作为稀疏帧推理的参考。在强化后训练期间,配对的密集和稀疏视图通过 grounding 奖励和可靠性门控参考奖励进行优化,鼓励稀疏视图预测保留与任务相关的时间证据。值得注意的是,SAVER仅在1250个随机采样的时间 grounding 示例上进行训练,未使用任何视频问答注释。在三个时间 grounding 基准和六个视频问答基准上,SAVER在不同帧预算下始终提升性能。特别是,SAVER在使用显著更少帧的情况下,可达到或超越密集帧的Qwen3.5基线。这些结果表明,时间 grounding 可作为学习稀疏视频推理的有效证据定位代理,该推理可迁移至更广泛的视频理解任务。
英文摘要
Video-language models commonly assume that more temporal observations lead to more reliable reasoning. We question this assumption and argue that the key challenge is not merely processing more video frames efficiently, but learning to reason reliably under limited temporal evidence. We propose SAVER, a dense-to-sparse post-training framework that uses dense video views as training-time references for sparse-frame inference. During reinforcement post-training, paired dense and sparse views are optimized with grounding rewards and a reliability-gated reference reward, encouraging sparse view predictions to preserve task-relevant temporal evidence. Notably, SAVER is trained only on 1,250 randomly sampled temporal grounding examples, without using any video question answering annotations. Across three temporal grounding benchmarks and six video question-answering benchmarks, SAVER consistently improves performance across frame budgets. In particular, SAVER can match or surpass dense-frame Qwen3.5 baselines while using substantially fewer frames. These results show that temporal grounding can serve as an effective evidence-localization proxy for learning sparse video reasoning that transfers to broader video understanding tasks.