arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Spheriverse:野外球面观测下的三维场景理解

Spheriverse: 3D Scene Understanding from Spherical Observations in the Wild

Fei Teng, Sheng Wu, Mengfei Duan, Guoqiang Zhao, Junhui Ma, Kai Luo, Siyu Li, Hao Shi, Zhiyong Li, Kailun Yang

arXiv 2609.09012首次发表:更新:

发表机构

Hunan University; Zhejiang University of Science and Technology; Ant Group; Zhejiang University(湖南大学; 浙江科技学院; 蚂蚁集团; 浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Spheriverse提出球面观测下的三维场景理解方法,含大规模数据集和SphereOcc框架,通过球面几何建模与证据检索,在占用预测等任务上超越现有方法。

AI 中文摘要

球面观测为三维场景理解提供了全局视觉上下文。然而,视觉信息在角度域中编码,而物理世界以笛卡尔坐标表示。这种跨空间表示差距使几何对应和语义证据聚合变得复杂。为深入探究这一挑战,我们引入了Spheriverse,包含64,400对时间对齐的球面图像-激光雷达数据对,组织成644个序列。该数据集涵盖多样的场景、光照和天气条件,并具有细粒度的语义类别。我们进一步建立了语义占用预测、语义建图和三维目标检测的基准,通过整体和场景级比较评估了30多种方法。对于密集预测,我们提出了SphereOcc,一种将球面几何建模与语义证据检索相结合的占用框架。笛卡尔-球面表示重塑(CSRR)通过区域调制将球面距离-方位角几何融入笛卡尔体素特征。球面证据重查询(SER)随后根据体素内容和距离-高度-方位角几何对查询进行条件化,以自适应地从源球面图像特征中检索相关语义证据。SphereOcc实现了13.91%的mIoU和24.65%的GeoIoU,分别比最优方法TPVFormer和SurroundOcc高出1.70和2.10个百分点。它还在所有五个场景的这两项指标中均排名第一,在评估的空间分区和减小的视场中具有一致的优势。所建立的基准和源代码将在此https URL上提供。

英文摘要

Spherical observations provide global visual context for 3D scene understanding. However, visual information is encoded in an angular domain, whereas the physical world is represented in Cartesian coordinates. This cross-space representation gap complicates geometric correspondence and semantic evidence aggregation. To delve into this challenge, we introduce Spheriverse, comprising 64,400 temporally aligned spherical image-LiDAR pairs organized into 644 sequences. The dataset spans diverse scenes, illumination, and weather conditions, with fine-grained semantic classes. We further establish benchmarks for semantic occupancy prediction, semantic mapping, and 3D object detection, evaluating 30+ methods through overall and scene-wise comparisons. For dense prediction, we propose SphereOcc, an occupancy framework that couples spherical geometry modeling with semantic evidence retrieval. Cartesian-Spherical Representation Remodeling (CSRR) incorporates spherical range-azimuth geometry into Cartesian voxel features through region-wise modulation. Spherical Evidence Re-querying (SER) then conditions queries on voxel content and range-height-azimuth geometry to adaptively retrieve relevant semantic evidence from source spherical image features. SphereOcc achieves 13.91% mIoU and 24.65% GeoIoU, yielding relative improvements of 13.9% and 9.3% over the respective best-performing methods, TPVFormer and SurroundOcc. It also ranks first in both metrics across all five scene categories, with consistent advantages across the evaluated spatial partitions and reduced fields of view. The established benchmark and source code will be available at https://feit-feiteng.github.io/Spheriverse.

CommentsThe established benchmark and source code will be available at https://feit-feiteng.github.io/Spheriverse

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑