发表机构
University of Windsor(温莎大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出GPT和ASH两项技术,将SAM3扩展为零样本开放词汇多目标跟踪与分割,实现任意长度视频的自动标注,在MOTS20上达到最先进性能且内存低于25GB。
AI 中文摘要
基于记忆注意力的视频实例分割(VIS)方法已展现出强大的零样本跟踪能力,然而其庞大的内存需求将其限制于短视频片段,且其单提示推理设计使得多类别开放词汇跟踪在计算上难以承受。本工作为任意视频的全自动跟踪标注引入两项贡献。广义存在令牌(GPT)通过虚拟提示批处理重构SAM3的推理流程,以同时处理N个文本提示,将图像编码成本从O(N)降至O(1),且无需修改任何已学习组件。标注与分割处理器(ASH)通过重叠时间块及基于IoU的块间身份匹配,将任何记忆注意力VIS跟踪器扩展至任意长度序列,无需特定数据集训练。实例化于SAM3上,所得流程——SAM3-ASH——在完全零样本条件下于MOTS20上达到最先进的HOTA,并在另外七个基准上与经过训练的专家保持竞争力,同时峰值GPU内存消耗保持在25 GB以下,为可扩展、免训练的全自动视频标注建立了实用基线。
英文摘要
Memory-attention-based Video Instance Segmentation (VIS) methods have demonstrated strong zero-shot tracking capability, yet their substantial memory requirements confine them to short video clips and their single-prompt inference design makes multi-category open-vocabulary tracking computationally prohibitive. This work introduces two contributions toward fully automated tracking annotation of arbitrary video. The Generalized Presence Token (GPT) reformulates SAM3's inference pipeline to process N text prompts simultaneously via virtual prompt batching, reducing image encoding cost from O(N) to O(1) with no modifications to any learned component. The Annotation and Segmentation Handler (ASH) extends any memory-attention VIS tracker to sequences of arbitrary length through overlapping temporal chunks with IoU-based inter-chunk identity matching, requiring no dataset-specific training. Instantiated on SAM3, the resulting pipeline -- SAM3-ASH -- achieves state-of-the-art HOTA on MOTS20 under fully zero-shot conditions and remains competitive with trained specialists across seven additional benchmarks, while peak GPU memory consumption stays below 25 GB, establishing a practical baseline for scalable, training-free automated video annotation.