arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.22687cs.CV

PanoSeg3R:基于自动数据策展流程的全景图像前馈式3D语义分割

PanoSeg3R: Feed-Forward 3D Semantic Segmentation for Panoramic Images with an Automatic Data Curation Pipeline

Heechan Yoon, Dongki Jung, Phuc Nguyen, Ming Lin, Dinesh Manocha

首次发表
浏览论文内容

中文总结 AI 辅助

PanoSeg3R提出前馈式3D全景语义分割框架,联合预测几何与多视角语义,并引入自动数据策展流程,显著提升零样本泛化性能。

中文摘要 AI 辅助

我们提出了PanoSeg3R,一种用于3D全景语义分割的前馈式框架。与现有针对透视输入设计的方法不同,PanoSeg3R在单次前向传播中联合预测3D几何和多视角语义分割。基于支持全景图像的预训练重建骨干网络,我们的方法通过查询式掩码解码器扩展了前馈式3D重建。此外,我们引入了一个自动全景数据策展流程,利用现成基础模型的互补优势生成可靠的伪语义标注,大幅扩展训练数据并提升零样本泛化能力。PanoSeg3R在全景3D语义分割上达到了最先进的性能,在ScanNet++上将3D mIoU提升了高达16.02,而策展后的训练数据进一步将Stanford2D3D和ToF-360上的零样本性能分别提升了高达4.26和43.28 mIoU。网站:此https URL

英文摘要

We present PanoSeg3R, a feed-forward framework for 3D panoramic semantic segmentation. Unlike existing methods designed for perspective inputs, PanoSeg3R jointly predicts 3D geometry and multi-view semantic segmentation in one single forward pass. Built upon a pretrained reconstruction backbone that supports panoramic images, our approach extends feed-forward 3D reconstruction with a query-based mask decoder. Furthermore, we introduce an automatic panorama data curation pipeline that leverages the complementary strengths of off-the-shelf foundation models to generate reliable pseudo semantic annotations, substantially expanding the training data and improving zero-shot generalization. PanoSeg3R achieves state-of-the-art performance on panoramic 3D semantic segmentation, improving 3D mIoU by up to 16.02 on ScanNet++, while the curated training data further improves zero-shot performance by up to 4.26 and 43.28 mIoU on Stanford2D3D and ToF-360, respectively. Website: https://harryyoon777.github.io/PanoSeg3R/

发表机构

  • University of Maryland, College Park(马里兰大学学院公园分校)

机构由 AI 辅助整理,请以论文原文为准。

↑