arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34895cs.CV

LVMT:用于长时视频分割的视频掩码Transformer

LVMT: Video Mask Transformer for Long-term Video Segmentation

Narges Norouzi, Niccolò Cavagnero, Idil Esen Zulfikar, Bastian Leibe, Gijs Dubbelman, Daan de Geus

首次发表
浏览论文内容

中文总结 AI 辅助

针对在线视频分割在长视频中长期遮挡下跟踪困难的问题,提出LVMT模型,采用轻量级GRU时间传播模块和截断查询传播训练策略,在六个基准上达到最先进性能,速度提升10倍。

中文摘要 AI 辅助

现有的在线视频分割方法在处理具有长期遮挡的长且复杂的视频时,难以跟踪目标。我们假设这一局限性源于:(i)其时间传播机制无法自适应地选择跨时间传播的目标信息,以及(ii)由于内存需求和梯度消失问题,它们无法在长视频上进行训练。为解决第一个局限性,我们提出使用一个轻量级的基于GRU的时间传播模块,该模块能够学习选择在记忆中保留哪些信息并跨时间传播。其次,为了允许在长视频上进行训练,我们引入了截断查询传播(Truncated Query Propagation, TQP)训练策略,在该策略中,模型按帧块处理视频,跟踪目标的信息在块之间传播,但反向传播仅在各个块内进行,从而在不出现内存溢出、推理开销或梯度消失的情况下实现更长的时间监督。由此产生的模型被称为长时视频掩码Transformer(Long-term Video Mask Transformer, LVMT)。在六个基准上的大量实验表明,LVMT在一系列视频分割任务上达到了新的最先进水平,同时保持了其基于的高效模型的速度,比先前的最先进方法快10倍。代码:此https URL

英文摘要

Existing online video segmentation methods struggle to track objects in long, complex videos with long-term occlusions. We hypothesize that this limitation is caused by (i) the inability of their temporal propagation mechanism to adaptively select the object information that is propagated across time, and (ii) their inability to be trained on long videos due to memory requirements and vanishing gradients. To address the first limitation, we propose to use a lightweight GRU-based temporal propagation module that can learn to select which information it keeps in memory and propagates across time. Second, to allow training on long videos, we introduce Truncated Query Propagation (TQP), a training strategy in which the model processes a video in chunks of frames, where information about tracked objects is propagated between chunks but backpropagation is only conducted in individual chunks, enabling longer temporal supervision without out-of-memory issues, inference overhead, or vanishing gradients. The resulting model is called the Long-term Video Mask Transformer (LVMT). Extensive experiments on six benchmarks show that LVMT sets a new state of the art across a range of video segmentation tasks, while retaining the speed of the highly efficient model it is based on, making it 10X faster than the prior state of the art. Code: https://www.tue-mps.org/lvmt

发表机构

  • Eindhoven University of Technology(埃因霍温理工大学)
  • RWTH Aachen University(亚琛工业大学)

机构由 AI 辅助整理,请以论文原文为准。

↑