arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

长到短视频证据推理用于接地问答

Long-to-Short Video Evidence Reasoning for Grounded Question Answering

Kaiyan Chen, Junbin Xiao, Xun Yang

arXiv 2609.15224首次发表:更新:

发表机构

University of Science and Technology of China(中国科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

LOVER通过长到短证据课程学习、IoP奖励和自适应时间戳渲染三项创新,在NExT-GQA和ReXTime基准上达到开源模型SOTA,有效提升接地问答性能。

AI 中文摘要

我们提出了LOVER,一个长到短视频证据强化模型,用于接地问答(GQA)。与现有的基于强化学习(RL)的视频推理模型相比,LOVER突出了三项创新:(1)长到短视频证据课程学习,它根据证据持续时间组织RL训练,并逐步使模型从长程接地适应到短程推理;(2)GQA奖励,强调IoP奖励在证据定位方面优于IoU,而非严格的时序跨度重叠;(3)自适应时间戳渲染,利用背景感知的位置和颜色选择,将时间戳自适应地渲染到视频帧上,以增强时间可观察性。这三个设计是模型无关且相互促进的。它们有效提升了不同骨干网络上的问答、接地以及接地问答性能。值得注意的是,基于Time-R1构建的LOVER在流行的GQA基准NExT-GQA和ReXTime上,达到了开源模型中的最新技术水平(SOTA)。全面的消融研究进一步验证了我们三个创新组件的有效性。

英文摘要

We present LOVER, a \underline{L}ong to sh\underline{O}rt \underline{V}ideo \underline{E}vidence \underline{R}einforced model for grounded question answering (GQA). LOVER highlights three innovations over existing reinforcement-learning (RL) based video reasoning models: (1) \textbf{Long-to-short Video Evidence Curriculum Learning}, which organizes RL training according to evidence duration and progressively adapts the model from long-range grounding to short-term reasoning; (2) \textbf{GQA Rewards}, which underscore the benefit of IoP reward over IoU for evidence spotting rather than strict temporal span overlap; (3) \textbf{Adaptive Timestamp Rendering}, which adaptively renders timestamps onto video frames using background-aware position and color selection to enhance temporal observability. The three designs are model-agnostic and reciprocal. They effectively improve QA, grounding, and grounded QA performance over different backbones. Notably, LOVER built on Time-R1 achieves new state-of-the-art (SOTA) results among open-source models on popular GQA benchmarks: NExT-GQA and ReXTime. Comprehensive ablation studies further validate the effectiveness of our three innovative components.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑