VOS-Agent:第8届LSVOS挑战赛(MOSEv2赛道)的第一名解决方案
VOS-Agent: The 1st Place Solution for the 8th LSVOS Challenge (MOSEv2 Track)
- Harbin Institute of Technology, Shenzhen(哈尔滨工业大学(深圳))
- Nanyang Technological University(南洋理工大学)
- Shenzhen Loop Area Institute(深圳河套学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对复杂视频目标分割中微小目标、语义主导目标的稳健传播难题,本文提出VOS-Agent协作框架,以SAM3为共享密集分割模块,根据目标特征激活专用智能体,在ECCV 2026第8届LSVOS挑战赛MOSEv2赛道获第一名,J&JF指标达69.82%。
AI中文摘要:
复杂视频目标分割需要在严重遮挡、目标消失与重现的情况下实现稳健的目标传播。尽管SAM3提供了强大的可提示掩码传播能力,但统一推理路径对于视觉证据不足的微小目标、以及身份依赖显式属性的语义主导目标而言仍不可靠。为此,我们提出VOS-Agent,这是一个保留SAM3作为共享密集分割模块、并根据目标特征有条件激活专用智能体的协作框架。目标感知与路由智能体将每个序列分配至常规、微小或语义主导的路由;微小目标由视觉跟踪智能体通过置信度感知的框提示提供支持,语义主导目标则由基于MLLM的语义智能体通过描述引导的定位与候选验证处理。在MOSEv2测试集上,VOS-Agent在官方$\boldsymbol{\textit{J}\boldsymbol{\textit{&}}\boldsymbol{\textit{J}}\boldsymbol{\textit{F}}$指标上取得69.82%的成绩,在ECCV 2026举办的第8届LSVOS挑战赛MOSEv2赛道中排名第一。
英文摘要:
Complex video object segmentation requires robust target propagation under severe occlusion, disappearance and reappearance. Although SAM3 provides strong promptable mask propagation, a uniform inference path remains unreliable for tiny targets with insufficient visual evidence and semantic-dominated targets whose identities depend on explicit attributes. To this end, we present VOS-Agent, a collaborative framework that retains SAM3 as the shared dense segmentation module and conditionally activates specialized agents according to target characteristics. A Target Perception and Routing Agent assigns each sequence to a regular, tiny, or semantic-dominated route. Tiny targets are supported by a Visual Tracking Agent through confidence-aware box prompts, while semantic-dominated targets are handled by an MLLM-based Semantic Agent through description-guided localization and candidate verification. On the MOSEv2 test set, VOS-Agent achieves 69.82% on the official $\mathcal{J}\&\dot{\mathcal{F}}$ metric and ranks first in the MOSEv2 Track of the 8th LSVOS Challenge at ECCV 2026.