arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31725cs.CVcs.AI

SWT:基于滑动窗口、小波变换和最优传输的自监督视频目标分割

SWT: Self-Supervised Video Object Segmentation via Sliding, Wavelet and Transportation

Zhengtong Zhu, Jiaqing Fan, Hanwen Qian, Fanzhang Li

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出SWT,一种基于滑动窗口、小波变换和最优传输的自监督视频目标分割框架,仅在COCO上训练一次,在五个VOS数据集上取得优异结果。

中文摘要 AI 辅助

视频目标分割(VOS)旨在从连续视频帧中准确分割目标对象,并跟踪对象在视频每一帧中的变化。传统的VOS方法通常需要大量像素级标注的视频序列进行全监督学习,这限制了模型在稀疏视频场景中的性能,而现有的VOS方法对对象的全局变化适应性有限。基于这一观察,本文提出了自监督VOS方法——滑动窗口、小波变换和最优传输(SWT),这是一种完全在静态数据集上使用对比学习训练的自监督VOS框架。首先,滚动样本缓冲区在连续更新中重用独立采样图像的重复组。其次,为了解决简单卷积结构造成的远距离建模困难,我们引入小波变换来扩大卷积核的感受野,从而提高模型的表示能力。最后,我们结合最优传输帮助模型在目标跨两帧之间找到全局最优匹配,提高模型处理对象非刚性变形的能力。SWT仅在COCO数据集上训练一次,就在五个VOS数据集以及一个额外的身体部位传播数据集上取得了优异的结果。代码将在[此HTTPS链接](此HTTPS链接)上很快发布。

英文摘要

Video Object Segmentation (VOS) aims to accurately segment target objects from consecutive video frames and track the changes of the objects in each frame of the video. Conventional VOS methods typically demand substantial quantities of pixel-level labeled video sequences for fully supervised learning, which limits the performance of the model in sparse video scenes, while existing VOS methods have limited adaptability to global changes in objects. Based on this observation, in this paper, we propose self-supervised VOS with Sliding window, Wavelet transform and optimal Transport (SWT), a self-supervised VOS framework entirely trained on static dataset using contrastive learning. Firstly, a rolling sample buffer reuses overlapping groups of independently sampled images across successive updates. Secondly, to address the long-distance modeling difficulty caused by simple convolutional structures, we introduce wavelet transform to expand the receptive field of convolutional kernels, thus improving the model's representational capability. Finally, we incorporate optimal transport to assist the model in finding the globally optimal match between the target across two frames, improving the model's ability to handle nonrigid deformations of objects. SWT only requires training on the COCO dataset once and achieves excellent results on five VOS datasets as well as an additional body part propagation dataset. The code will be released soon at [https://github.com/machine928/SWT.git](https://github.com/machine928/SWT.git).

↑