AdaOcc:面向具身任务的自适应三维占据预测
AdaOcc: Adaptive 3D Occupancy Prediction for Embodied Tasks
浏览论文内容
中文总结 AI 辅助
本文提出AdaOcc,一种基于点的自适应三维占据预测方法,通过几何引导双分支编码器和包含损失,在Occ-ScanNet上达到新最优,并支持灵活计算预算与异构传感器输入。
中文摘要 AI 辅助
具身任务要求准确、灵活且语义丰富的三维场景表示。三维语义占据通过编码几何占据和语义类别来建模整体三维空间,非常适合这一需求。然而,现有的占据预测方法难以满足实际部署要求,如适应不同的计算预算、传感器配置和观察视角。本文提出了一种基于点的自适应三维占据预测方法,称为AdaOcc,专为具身场景设计。为适应异构传感器输入,AdaOcc采用自适应几何引导的双分支编码器,可支持不同视角数量的RGB图像以及(估计的)深度图或激光雷达扫描。AdaOcc通过渐进式查询学习策略用稀疏语义点表示占据区域,允许通过查询点数量和解码器层数灵活调整预测计算预算。为促进轻量级基于点的占据学习中的高保真几何建模,我们进一步提出了一种新颖的包含损失,该损失将预测点正则化到有效占据区域内。大量实验表明,我们的方法在Occ-ScanNet上达到了新的最先进水平,相比之前的方法有显著性能提升。此外,我们的框架作为自适应三维感知模块在真实世界具身系统中展示了强大的实际适用性。
英文摘要
Embodied tasks demand accurate, flexible, and semantically rich 3D scene representations. 3D semantic occupancy is well suited to this requirement, as it can model holistic 3D spaces by encoding geometric occupancy along with semantic categories. However, existing occupancy prediction methods struggle to meet practical deployment requirements, such as adapting to varying computing budgets, sensor setups, and observation views. In this paper, we propose a point-based Adaptive 3D Occupancy Prediction method, called AdaOcc, tailored for embodied scenarios. To accommodate heterogeneous sensor inputs, AdaOcc uses an adaptive geometry-guided dual-branch encoder that can support RGB images in various numbers of views with (estimated) depth maps or LiDAR scans. AdaOcc represents occupied regions via sparse semantic points trained with a progressive query learning strategy, allowing the prediction computational budget to be flexibly adjusted through query point numbers and decoder layers. To facilitate high-fidelity geometric modeling for lightweight point-based occupancy learning, we further propose a novel containment loss that regularizes predicted points to reside within valid occupied regions. Extensive experiments show that our method achieves a new state-of-the-art on Occ-ScanNet with considerable performance improvements over previous methods. Moreover, our framework demonstrates strong practical applicability as an adaptive 3D perception module in real-world embodied systems.
发表机构
- Beihang University(北京航空航天大学)
- Beijing Academy of Artificial Intelligence(北京人工智能研究院)
- XYZ Embodied AI(XYZ具身智能公司)
- Hebei University of Technology(河北工业大学)
- ShanghaiTech University(上海科技大学)
- University of Science and Technology Beijing(北京科技大学)
机构由 AI 辅助整理,请以论文原文为准。