面向视频以对象为中心学习的语义槽
Semantic Slots for Video Object-Centric Learning
查看机构详情
- Polytechnique Montréal(蒙特利尔理工学院)
- Université TÉLUQ(魁北克大学泰吕克分校)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
针对视频以对象为中心学习的解码器瓶颈,提出SemanticSlots方法,通过Transformer解码器使槽与对象位置无关,在YouTube-VIS数据集上实现优于现有方法的分割性能。
中文摘要 AI 辅助
视频以对象为中心学习(OCL)传统上聚焦于优化编码器架构以确保时间一致性。本文指出,主要瓶颈在于解码器:传统解码器强制槽(slots)空间锚定,阻碍其适应运动的能力。我们提出SemanticSlots,采用基于Transformer的解码器,利用图像上下文,使槽无需编码边界精度和空间位置,可作为本质上与对象位置无关的语义查询,检索匹配特征而非记忆坐标。更重要的是,该特性允许从单帧计算的槽分解后续视频帧,无需复杂的时间预测器或辅助时间损失。在YouTube-VIS数据集上,SemanticSlots的mBO指标较VideoSAUR提升31个点,较当前SOTA方法提升21个点,达到86.6%的ARI和62.8%的mBO。
英文摘要
Video Object-Centric Learning (OCL) has traditionally focused on refining the encoder architecture to ensure temporal consistency. In this paper, we argue that the primary bottleneck lies in the decoder. We show that traditional decoders force slots to be spatially anchored, hindering their ability to adapt to motion. We propose SemanticSlots, which uses a Transformer-based decoder that leverages image context, relieving slots from encoding boundary precision and spatial location. This allows slots to function as semantic queries that are inherently object position invariant, retrieving matching features rather than memorizing coordinates. More importantly, this property allows slots computed from a single frame to decompose subsequent video frames, eliminating the need for complex temporal predictors or auxiliary temporal losses. Results on YouTube-VIS show that SemanticSlots improves upon VideoSAUR by 31 points in mBO and outperforms current state-of-the-art methods by 21 points, achieving 86.6% ARI and 62.8% mBO.