arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

每一视图都重要:多视图全景分割的视图一致全景质量

Every View Counts: View-Consistent Panoptic Quality for Multi-view Panoptic Segmentation

Youngmin Lee, Byungha Ko, Guhnoo Yun, Dong Hwan Kim

arXiv 2610.05911首次发表:更新:

发表机构

Korea University; Korea Institute of Science and Technology(高丽大学; 韩国科学技术研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出VC-PQ指标,将全景分割评估从单图像扩展到多视图,均等计数每个视图,惩罚视图不一致预测,分解损失来源,并在ScanNet++和ScanNetv2上验证其优于PQ^scene的敏感性。

AI 中文摘要

多视图全景分割为无序图像集中的每个像素分配一个语义类别和一个场景级实例ID,最近的feed-forward 3D模型在单次前向传播中为输入视图预测这些标签。然而,它们的预测一直使用从逐场景优化方法借用的场景级PQ(PQ^scene)进行评估,通常在渲染的保留视图上进行。PQ^scene将场景的所有视图拼接成一张图像,因此遗漏的外观或ID变化只会按面积比例降低匹配对的分数。我们提出视图一致全景质量(VC-PQ),它将PQ从单张图像扩展到一组输入视图,对实例可见的每个视图均等计数,并惩罚在与其真值相同的视图中不可见的预测。VC-PQ的分解将方法损失的分数归因于掩码精度、视图一致性和匹配阈值。一个额外的参数可恢复拼接的面积加权以进行比较。在ScanNet++和ScanNetv2上的固定评估协议下,使用VC-PQ和PQ^scene评估了最近的feed-forward方法,分解显示了每种方法在何处损失分数。对真值的受控扰动表明,VC-PQ响应于实例被遗漏或ID变化的视图数量,而PQ^scene响应于其面积。这项工作的目的是使视图一致性成为多视图全景分割评估的一部分,将VC-PQ与PQ^scene一起报告。

英文摘要

Multi-view panoptic segmentation assigns a semantic class and a scene-level instance ID to every pixel of an unordered set of images, and recent feed-forward 3D models predict these labels for the input views in a single forward pass. Their predictions, however, have been evaluated with the scene-level PQ (PQ^scene) borrowed from per-scene optimization methods, typically on rendered held-out views. PQ^scene tiles all views of a scene into a single image, so that a missed appearance or a change of ID lowers the score of the matched pair only in proportion to its area. We propose View-Consistent Panoptic Quality (VC-PQ), which extends PQ from a single image to a set of input views, counts equally every view in which an instance is visible, and penalizes a prediction that is not visible in the same views as its ground truth. A decomposition of VC-PQ attributes the score a method loses to mask accuracy, view consistency, and the matching threshold. A single additional parameter recovers the area weighting of tiling for comparison. Under a fixed evaluation protocol on ScanNet++ and ScanNetv2, recent feed-forward methods are evaluated with VC-PQ and PQ^scene, and the decomposition shows where each of them loses its score. Controlled perturbations of the ground truth show that VC-PQ responds to the number of views in which an instance is missed or changes ID, whereas PQ^scene responds to their area. The aim of this work is to make view consistency part of the evaluation of multi-view panoptic segmentation, with VC-PQ reported alongside PQ^scene.

Comments23 pages, 8 figures. Under review. Youngmin Lee and Byungha Ko contributed equally

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑