arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在视频中定位任意目标:重新思考高效的生成式时空视频定位

Locate Anything in Videos: Rethinking Efficient Generative Spatio-Temporal Video Grounding

Hanoona Rasheed, Haania Siddiqui, Ming-Hsuan Yang, Fahad Shahbaz Khan, Salman Khan

arXiv 2608.28192首次发表:更新:

发表机构

Mohamed bin Zayed University of Artificial Intelligence; University of California, Merced; Apertix(穆罕默德·本·扎耶德人工智能大学; 加州大学默塞德分校; 阿珀蒂克斯公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对时空视频定位的自回归解码存在延迟高、误差易传播的问题,提出并行轨迹解码,在VidSTG等数据集上实现显著的延迟降低与吞吐量提升,且可零样本泛化到多项相关任务

AI 中文摘要

时空视频定位(STVG)要求模型识别所提及事件发生的时间,并在该时间区间内定位目标实体。现有的多模态大语言模型通常以自回归方式序列化密集的定位轨迹,导致解码延迟随轨迹长度增长,且定位误差会跨时间传播。我们提出并行轨迹解码(PTD),这是一种生成式范式,将定位分解为时间块和后续的时间条件空间块,二者可同时解码。这消除了令牌级和轨迹级依赖,将顺序解码深度降低为固定的1+1轮,与轨迹长度无关。为实现并行空间生成,我们引入了解耦块注意力,它在保留对共享视频-查询上下文访问权限的同时消除了框间依赖,还结合了针对时间边界和空间几何的感知定位策略优化。在VidSTG数据集上,与标准自回归解码相比,PTD将轨迹完成延迟降低了79倍,空间解码吞吐量提高了92倍,同时还提升了定位精度。采用紧凑的4B骨干网络,我们的模型在VidSTG和HC-STVG上表现出色,并且可零样本泛化到时间定位、基于定位的视频问答以及指代视频目标跟踪任务。我们的结果表明,并行轨迹生成是视频中自回归定位的一种高效且有效的替代方案。

英文摘要

Spatio-temporal video grounding (STVG) requires models to identify when a referred event occurs and localize the target entity throughout that interval. Existing multimodal large language models typically serialize dense localization trajectories autoregressively, causing decoding latency to grow with tube length and allowing localization errors to propagate across time. We introduce Parallel Tube Decoding (PTD), a generative formulation that decomposes grounding into a temporal block followed by time-conditioned spatial blocks decoded simultaneously. This removes both token-level and trajectory-level dependencies, reducing the sequential decoding depth to a fixed $1 + 1$ rounds, independent of tube length. To enable parallel spatial generation, we introduce Decoupled Block Attention, which preserves access to shared video-query context while eliminating cross-box dependencies, together with localization-aware policy optimization for temporal boundaries and spatial geometry. On VidSTG, PTD reduces Tube Completion Latency by 79x and increases spatial decoding throughput by 92x over standard autoregressive decoding, while also improving grounding accuracy. With a compact 4B backbone, our model performs favorably well on VidSTG and HC-STVG, and generalizes zero-shot to temporal grounding, grounded VideoQA, and referring video object tracking. Our results show parallel tube generation is an efficient and effective alternative to autoregressive localization in videos.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑