arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MoVISA:视频对象分割的多令牌推理

MoVISA: Multi-Token Reasoning for Video Object Segmentation

Ruining Zhao, Ho Kei Cheng, Alexander G Schwing

arXiv 2609.28956首次发表:更新:

发表机构

University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

MoVISA提出多令牌推理方法,通过多个分割令牌实现视频对象分割中语言提示与时空掩码的精细对齐,在MeViS和ReVOS上分别提升13.2%和8.4%的J和F指标。

AI 中文摘要

近年来,利用多模态大语言模型(MLLM)推理进行视频对象分割的研究进展表明,使用单个文本令牌(如SEG)来预测跨图像和视频的分割掩码是有效的。然而,我们观察到这种单令牌策略缺乏在视频分割任务中跨时间精确定位多个对象所需的粒度。为解决这一局限,我们开发了视频对象分割的多令牌推理方法,即MoVISA。MoVISA使用多个分割令牌(如SEG0和SEG1)来表示不同帧中的对象。这种设计使得语言提示与时空掩码预测之间的对齐更加精细,从而提升了性能和可解释性。在具有挑战性的MeViS、DAVIS17、ReVOS和Ref-Youtube-VOS基准上,我们的模型在MeViS上实现了13.2%的J和F提升,在ReVOS上实现了8.4%的J和F提升。代码和模型将发布。

英文摘要

Recent advances in video object segmentation with Multimodal Large Language Model (MLLM) reasoning have demonstrated the effectiveness of using a single textual token, such as SEG, to predict segmentation masks across images and videos. However, we observe that this single-token strategy lacks the granularity required to precisely localize multiple objects across time in video segmentation tasks. To address this limitation, we develop Multi-Token Reasoning for Video Object Segmentation, or MoVISA. MoVISA uses multiple segmentation tokens, such as SEG0 and SEG1, to represent an object across different frames. This design enables more fine-grained alignment between language prompts and spatio-temporal mask predictions, improving both performance and interpretability. On the challenging MeViS, DAVIS17, ReVOS, and Ref-Youtube-VOS benchmarks, our model achieves a 13.2 percent J and F improvement on MeViS and an 8.4 percent J and F improvement on ReVOS. Code and models will be released.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑