硬区域监督:Waymo 开放数据集2D视频全景分割排行榜第一名
Hard-Region Supervision: #1 on the Waymo Open Dataset 2D Video Panoptic Segmentation Leaderboard
浏览论文内容
中文总结 AI 辅助
本文提出硬区域监督(HRS)方法,通过辅助预测头强化基线错误区域,并采用集成、输出合并和跨相机链接,在Waymo视频全景分割挑战中三项指标均获第一。
中文摘要 AI 辅助
我们描述了在Waymo开放数据集2D视频全景分割挑战赛中的获胜方案。该任务要求为每一帧的每个像素赋予一个语义类别,并且对于可计数对象,赋予一个在100帧和五个重叠相机之间保持一致的标识。我们以DVIS++(一个由分割器、跟踪器和精化器组成的级联模型)作为基线,并提出了硬区域监督(HRS)来改进基线。具体而言,我们利用基线定义硬区域为它出错的地方,并为该区域设计了一个损失函数和一个辅助预测头。该辅助头仅在训练中使用,在测试时被移除,因此使用HRS训练的模型在推理时与基线具有相同的架构。此外,我们提出了三个测试时步骤以进一步提升结果:双模型集成、将分割器的输出合并到最终全景图中,以及跨相机身份链接。在挑战测试集上,我们的方案达到了0.3547的wSTQ、0.2071的wAQ和0.6075的mIoU,在三个指标上均排名第一。它比第二名高出3.6个wSTQ点,比我们的DVIS++基线高出2.4个点。
英文摘要
We describe our winning entry to the Waymo Open Dataset 2D Video Panoptic Segmentation Challenge. The task asks for a semantic class at every pixel of every frame and, for countable objects, an identity that holds across 100 frames and across five overlapping cameras. We build on DVIS++, a cascade of a segmenter, a tracker, and a refiner, as our baseline. We propose hard region supervision (HRS) to improve the baseline. In particular, we use the baseline to define the hard region as where it makes mistakes, and design a loss and an auxiliary prediction head for this region. The auxiliary head is used only in training and removed at test time, so at inference the model trained with HRS has the same architecture as the baseline. In addition, we propose three test-time steps that further improve the results: a two-model ensemble, a merge of the segmenter's output into the final panoptic map, and cross-camera identity linking. On the challenge test set, our entry reaches 0.3547 wSTQ, 0.2071 wAQ, and 0.6075 mIoU, ranking first on all three metrics. It is 3.6 wSTQ points ahead of the second entry and 2.4 points ahead of our DVIS++ baseline.