arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从多中取少:利用密集参考强化稀疏视频推理

Less from More: Reinforcing Sparse Video Reasoning from Dense References

Wenfang Sun, Yingjun Du, Cees G. M. Snoek

arXiv 2610.10893首次发表:更新:

AI 中文总结

本研究质疑视频-语言模型依赖更多时间观测的假设,提出SAVER密集到稀疏后训练框架,仅用1250个时间grounding示例训练,可在少帧下提升稀疏视频推理性能,迁移至视频问答等任务。

AI 中文摘要

视频-语言模型通常假设更多的时间观测会带来更可靠的推理。我们质疑这一假设,并认为关键挑战不仅是高效处理更多视频帧,而是在有限的时间证据下学习可靠推理。我们提出SAVER,一种密集到稀疏的后训练框架,该框架在训练时使用密集视频视图作为稀疏帧推理的参考。在强化后训练期间,配对的密集和稀疏视图通过 grounding 奖励和可靠性门控参考奖励进行优化,鼓励稀疏视图预测保留与任务相关的时间证据。值得注意的是,SAVER仅在1250个随机采样的时间 grounding 示例上进行训练,未使用任何视频问答注释。在三个时间 grounding 基准和六个视频问答基准上,SAVER在不同帧预算下始终提升性能。特别是,SAVER在使用显著更少帧的情况下,可达到或超越密集帧的Qwen3.5基线。这些结果表明,时间 grounding 可作为学习稀疏视频推理的有效证据定位代理,该推理可迁移至更广泛的视频理解任务。

英文摘要

Video-language models commonly assume that more temporal observations lead to more reliable reasoning. We question this assumption and argue that the key challenge is not merely processing more video frames efficiently, but learning to reason reliably under limited temporal evidence. We propose SAVER, a dense-to-sparse post-training framework that uses dense video views as training-time references for sparse-frame inference. During reinforcement post-training, paired dense and sparse views are optimized with grounding rewards and a reliability-gated reference reward, encouraging sparse view predictions to preserve task-relevant temporal evidence. Notably, SAVER is trained only on 1,250 randomly sampled temporal grounding examples, without using any video question answering annotations. Across three temporal grounding benchmarks and six video question-answering benchmarks, SAVER consistently improves performance across frame budgets. In particular, SAVER can match or surpass dense-frame Qwen3.5 baselines while using substantially fewer frames. These results show that temporal grounding can serve as an effective evidence-localization proxy for learning sparse video reasoning that transfers to broader video understanding tasks.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑