arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

OmniAct3D:利用基础几何与证据接地推理实现全景3D检测

OmniAct3D: Leveraging Foundation Geometry and Evidence-Grounded Reasoning for Panoramic 3D Detection

Runtong Wu, Fei Teng, Di Wen, Guoqiang Zhao, Kunyu Peng, Kailun Yang

arXiv 2610.03015首次发表:更新:

发表机构

Hunan University; National Engineering Research Center of Robot Visual Perception and Control Technology; Karlsruhe Institute of Technology(湖南大学; 国家机器人视觉感知与控制技术工程研究中心; 卡尔斯鲁厄理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

OmniAct3D通过ERP射线几何适配器、视觉-动作推理链和外观引导朝向专家,将透视训练的VFM检测器适配到全景3D检测,显著提升NDS和mAP,并实现跨配置的可重用推理。

AI 中文摘要

精确的3D检测对于移动具身智能体至关重要,而视觉基础模型(VFM)提供了可迁移的视觉和几何先验。然而,现有的基于VFM的3D检测器依赖于窄视角的单目图像或离散的透视视图,限制了连贯的环绕感知;等距柱状投影(ERP)则能将连续的360度场景编码到单张图像中。直接迁移仍然困难,因为ERP以不同的方式组织几何和视觉信息,使得与物体相关的线索难以建模、定位和保留。我们提出了OmniAct3D,一个将透视训练的VFM检测器适配到ERP同时保留可迁移的VFM先验的框架。为解决几何不匹配问题,ERP射线几何适配器(ERGA-Ray)对球面视角射线和周期性空间结构进行建模。为在场景级上下文中定位证据,视觉-动作推理链(VARC)将每个假设基于相关的全景证据,并将其转换为结构化的几何动作。为恢复在固定令牌预算下丢失的局部线索,外观引导朝向专家(AGHE)以更高分辨率重新编码物体区域以进行朝向估计。实验表明,OmniAct3D在Spheriverse上比之前最佳的3D检测器提高了2.96个NDS点,在PanoMMOcc上比未适配的VFM基线提高了24.87个mAP点。通过目标特定的几何适配,VARC保持了相同配置mAP的95%至98%,表明物体级3D推理可跨传感配置重用。源代码将在提供的https URL上公开。

英文摘要

Accurate 3D detection is essential for mobile embodied agents, while Vision Foundation Models (VFMs) offer transferable visual and geometric priors. Yet existing VFM-based 3D detectors rely on narrow-view monocular images or discrete perspective views, limiting coherent surround perception; equirectangular projection (ERP) instead encodes a continuous 360 scene in a single image. Direct transfer remains difficult because ERP organizes geometry and visual information differently, making object-relevant cues hard to model, localize, and preserve. We propose OmniAct3D, a framework that adapts perspective-trained VFM detectors to ERP while preserving transferable VFM priors. To resolve geometric mismatch, the ERP-Ray Geometry Adapter (ERGA-Ray) models spherical viewing rays and periodic spatial structure. To localize evidence in scene-wide context, the Visual-Action Reasoning Chain (VARC) grounds each hypothesis in relevant panoramic evidence and converts it into a structured geometric action. To recover local cues lost under fixed token budgets, the Appearance-Guided Heading Expert (AGHE) re-encodes object regions at higher resolution for heading estimation. Experiments show that OmniAct3D improves over the previous best 3D detector by 2.96 NDS points on Spheriverse and over the unadapted VFM baseline by 24.87 mAP points on PanoMMOcc. With target-specific geometry adaptation, VARC retains 95--98% of the same-configuration mAP, indicating reusable object-level 3D reasoning across sensing configurations. The source code will be made publicly available at https://github.com/FeiT-FeiTeng/OmniAct3D.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑