发表机构
University of Science and Technology of China(中国科学技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
LOVER通过长到短证据课程学习、IoP奖励和自适应时间戳渲染三项创新,在NExT-GQA和ReXTime基准上达到开源模型SOTA,有效提升接地问答性能。
AI 中文摘要
我们提出了LOVER,一个长到短视频证据强化模型,用于接地问答(GQA)。与现有的基于强化学习(RL)的视频推理模型相比,LOVER突出了三项创新:(1)长到短视频证据课程学习,它根据证据持续时间组织RL训练,并逐步使模型从长程接地适应到短程推理;(2)GQA奖励,强调IoP奖励在证据定位方面优于IoU,而非严格的时序跨度重叠;(3)自适应时间戳渲染,利用背景感知的位置和颜色选择,将时间戳自适应地渲染到视频帧上,以增强时间可观察性。这三个设计是模型无关且相互促进的。它们有效提升了不同骨干网络上的问答、接地以及接地问答性能。值得注意的是,基于Time-R1构建的LOVER在流行的GQA基准NExT-GQA和ReXTime上,达到了开源模型中的最新技术水平(SOTA)。全面的消融研究进一步验证了我们三个创新组件的有效性。
英文摘要
We present LOVER, a \underline{L}ong to sh\underline{O}rt \underline{V}ideo \underline{E}vidence \underline{R}einforced model for grounded question answering (GQA). LOVER highlights three innovations over existing reinforcement-learning (RL) based video reasoning models: (1) \textbf{Long-to-short Video Evidence Curriculum Learning}, which organizes RL training according to evidence duration and progressively adapts the model from long-range grounding to short-term reasoning; (2) \textbf{GQA Rewards}, which underscore the benefit of IoP reward over IoU for evidence spotting rather than strict temporal span overlap; (3) \textbf{Adaptive Timestamp Rendering}, which adaptively renders timestamps onto video frames using background-aware position and color selection to enhance temporal observability. The three designs are model-agnostic and reciprocal. They effectively improve QA, grounding, and grounded QA performance over different backbones. Notably, LOVER built on Time-R1 achieves new state-of-the-art (SOTA) results among open-source models on popular GQA benchmarks: NExT-GQA and ReXTime. Comprehensive ablation studies further validate the effectiveness of our three innovative components.