arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34598cs.CVcs.AI

先总结后定位:查询引导的块压缩用于长视频时间定位

Summarize Before Grounding: Query-Guided Chunk Condensation for Long-Video Temporal Grounding

Nanxing Hu, Xiaoyue Duan, Qiwei Yan, Kailin Lyu, Jinchao Zhang, Guoliang Kang

首次发表
浏览论文内容

中文总结 AI 辅助

针对长视频时间定位中冗余信息干扰问题,提出SumGround框架,通过查询引导的块压缩和关联式摘要检索,结合RLVR训练与长度感知梯度门控,在多个数据集上超越现有方法。

中文摘要 AI 辅助

视频时间定位(VTG)旨在定位与语言查询相对应的视频片段。近期的大型视觉语言模型(LVLMs)在解决此类多模态推理任务方面展现出巨大潜力。然而,长视频通常包含大量冗余信息,这些信息会干扰LVLMs挖掘与查询相关的证据。不同于密集帧采样(其会导致难以承受的训练内存开销),以往基于可验证奖励的强化学习(RLVR)工作通常采用稀疏采样,这虽然使训练可行,但可能会遗漏关键证据。在本文中,我们提出了一种名为“SumGround”的“先总结后定位”框架,用于长视频时间定位。SumGround的关键在于执行查询引导的块压缩,以聚合和检索与查询相关的证据。具体而言,我们将视频分割成多个块,并执行两级块压缩。首先,我们引入查询引导的潜在摘要,其表示为查询引导提示的KV状态,以将冗余的视觉标记压缩为紧凑的、与查询相关的块摘要。此外,我们设计了一种关联式摘要检索方案,用于对最可能包含事件区间的块摘要进行排序和选择。查询引导的潜在摘要和关联式摘要检索方案均由RLVR实现。为减少内存消耗,我们提出了一种长度感知的梯度门控模块,以选择性地阻止梯度反向传播到视觉标记。大量实验表明,SumGround在多个下游数据集上优于先前的最先进方法,并在长视频上取得了显著提升。

英文摘要

Video temporal grounding (VTG) aims to localize the video interval corresponding to a language query. Recent large vision-language models (LVLMs) show great potential in solving such a multi-modal reasoning task. However, long videos often contain large amounts of redundant information that disturbs LVLMs to mine query-relevant evidence. Instead of dense frame sampling which incurs prohibitive training memory, previous reinforcement learning with verifiable rewards (RLVR) works typically utilize sparse sampling, which makes training feasible but may miss critical evidence. In this paper, we propose a ``summarize before grounding'' framework (named ``SumGround'') for long-video temporal grounding. The key of SumGround is to perform query-guided chunk condensation to aggregate and retrieve query-relevant evidence. Specifically, we split the video into several chunks and perform two-level chunk condensation. First, we introduce query-guided latent summaries, which is represented as KV states of query-guided prompts, to compress redundant visual tokens into compact query-relevant chunk summaries. Furthermore, we design an associative summary retrieval scheme to rank and select chunk summaries that are most likely to contain the event interval. Both query-guided latent summary and associative summary retrieval schemes are enabled by RLVR. To reduce memory consumption, we propose a length-aware gradient gating module to selectively stop gradient back-propagated to visual tokens. Extensive experiments demonstrate that SumGround performs favorably against previous state-of-the-art methods across multiple downstream datasets, with remarkable gains on long videos.

发表机构

  • Beihang University(北京航空航天大学)
  • University of the Chinese Academy of Sciences(中国科学院大学)
  • Institute of automation, Chinese academy of science(中国科学院自动化研究所)

机构由 AI 辅助整理,请以论文原文为准。

↑