第八届LSVOS挑战赛报告:复杂与多模态视频目标分割
Report of the 8th LSVOS Challenge: Complex and Multimodal Video Object Segmentation
浏览论文内容
中文总结 AI 辅助
本报告总结第八届LSVOS挑战赛,该赛在三个设置下评估视频分割,领先方案结合基础模型与多模态推理等模块,推动从单模型掩码传播向模块化流水线转变。
中文摘要 AI 辅助
本报告总结了与ECCV 2026同期举办的第八届大规模视频目标分割(LSVOS)挑战赛。该挑战赛在三个互补的设置下评估视频分割:在MOSEv2上的复杂半监督视频目标分割、在MeViSv2-Text上的文本引导的指代视频目标分割,以及在MeViSv2-Audio上的音频引导的指代视频目标分割。我们描述了任务和评估协议,并回顾了每个赛道前三名团队的方法。在九个领先的解决方案中,基础分割模型与目标感知记忆、多模态推理、显式目标存在性验证、智能体交互和修正性跟踪相结合。这些系统展示了从单模型掩码传播向模块化流水线的更广泛转变,这些流水线对目标身份、查询有效性和时间可靠性进行推理。
英文摘要
This report summarizes the 8th Large-scale Video Object Segmentation (LSVOS) Challenge, held in conjunction with ECCV 2026. The challenge evaluates video segmentation in three complementary settings: complex semi-supervised video object segmentation on MOSEv2, text-guided referring video object segmentation on MeViSv2-Text, and audio-guided referring video object segmentation on MeViSv2-Audio. We describe the tasks and evaluation protocols and review the methods of the top three teams in each track. Across the nine leading solutions, foundation segmentation models are combined with target-aware memory, multimodal reasoning, explicit target-existence verification, agentic interaction, and corrective tracking. These systems illustrate a broader transition from single-model mask propagation toward modular pipelines that reason about object identity, query validity, and temporal reliability.
发表机构
- Fudan University(复旦大学)
机构由 AI 辅助整理,请以论文原文为准。