感知干扰物的视频目标分割
Distractor-Aware Video Object Segmentation
- Linköping University(林雪平大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对视频目标分割中干扰物易引发误报的问题,提出感知干扰物的多对一方案,改进LWL方法,在DAVIS 2017数据集上实现最优性能并提升基线4.6个百分点。
AI中文摘要:
半监督视频目标分割是一项具有挑战性的任务,旨在给定第一帧的初始掩码,在整个视频序列中分割目标。判别式方法在合理的复杂度下已在该任务上展现出有竞争力的性能,这类方法通常将问题表述为目标与背景之间的一对一分类。然而,实际中视频序列通常包含目标、背景,还可能存在其他干扰物,这些物体增加了误报风险,尤其是当它们与目标具有视觉相似性时。因此,将干扰物与背景分离并独立处理会更有效。我们提出一种多对一方案,通过将干扰物划分为单独的类别来应对这种情况,该分离可让模型对最可能降低性能的挑战性区域施加特殊关注。我们通过修改学习学习(LWL)方法使其具备干扰物感知能力,证明了该方案的优越性。所提方法在DAVIS 2017验证数据集上达到了新的最优性能,在DAVIS 2017测试开发基准上较基线方法提升了4.6个百分点。
英文摘要:
Semi-supervised video object segmentation is a challenging task that aims to segment a target throughout a video sequence given an initial mask at the first frame. Discriminative approaches have demonstrated competitive performance on this task at a sensible complexity. These approaches typically formulate the problem as a one-versus-one classification between the target and the background. However, in reality, a video sequence usually encompasses a target, background, and possibly other distracting objects. Those objects increase the risk of introducing false positives, especially if they share visual similarities with the target. Therefore, it is more effective to separate distractors from the background, and handle them independently. We propose a one-versus-many scheme to address this situation by separating distractors into their own class. This separation allows imposing special attention to challenging regions that are most likely to degrade the performance. We demonstrate the prominence of this formulation by modifying the learning-what-to-learn (LWL) method to be distractor-aware. Our proposed approach sets a new state-of-the-art on the DAVIS 2017 val dataset, and improves over the baseline on the DAVIS 2017 test-dev benchmark by 4.6 percentage points.