用于鲁棒视频目标分割的竞争性记忆读出:第8届LSVOS挑战赛MOSEv2赛道第2名技术报告
Competitive Memory Readout for Robust Video Object Segmentation: 2nd Place Technical Report for the MOSEv2 Track of the 8th LSVOS Challenge
浏览论文内容
中文总结 AI 辅助
本研究针对第8届LSVOS挑战赛MOSEv2赛道提出基于SAM~3的改进方法,通过竞争性记忆读出机制提升鲁棒视频目标分割性能,获该赛道第2名。
中文摘要 AI 辅助
我们展示了针对ECCV 2026第8届大规模视频目标分割(LSVOS)挑战赛MOSEv2赛道的解决方案。该挑战赛评估复杂时序动态下的鲁棒视频目标分割能力,包括长期遮挡、消失与重现、外观大幅变化,以及来自视觉相似对象的强干扰。我们的方法基于SAM~3模型,聚焦其记忆读出环节。仅针对目标的标准记忆检索会将标注目标与同类非目标对象混淆,因为此类干扰项仅作为背景被隐式表示。我们的方法引入竞争性记忆读出机制,在从记忆中检索目标信息时明确纳入同类竞争对象的证据。为防止对弱目标或重现目标的过度抑制,我们在竞争后进一步应用轻量自适应恢复规则。所得系统保留了原始SAM~3跟踪流程,同时提升了挑战性视频中的目标身份保持能力。我们的提交在挑战赛主评分中取得66.20分,在MOSEv2赛道排名第2。
英文摘要
We present our solution for the MOSEv2 track of the 8th Large-scale Video Object Segmentation (LSVOS) Challenge at ECCV 2026. The challenge evaluates robust video object segmentation under complex temporal dynamics, including long-term occlusion, disappearance and reappearance, large appearance changes, and strong interference from visually similar objects. Our method builds on SAM~3 and focuses on its memory readout. Standard target-only memory retrieval can confuse the annotated target with same-class non-target objects because such distractors are represented only implicitly as background. Our method introduces Competitive Memory Readout, which explicitly incorporates same-class competitor evidence when retrieving target information from memory. To prevent excessive suppression of weak or reappearing targets, we further apply a lightweight adaptive restoration rule after competition. The resulting system retains the original SAM~3 tracking pipeline while improving target identity preservation in challenging videos. Our submission achieves 66.20 on the primary challenge score and ranks 2nd in the MOSEv2 track.
发表机构
- School of Computer Science, University of Sheffield(谢菲尔德大学计算机学院)
- Department of Automation, Tsinghua University(清华大学自动化系)
机构由 AI 辅助整理,请以论文原文为准。