Stream3Dv2:几何-语义融合增强的流式零样本3D场景理解
Stream3Dv2: Geometric-Semantic Fusion Enhanced Streaming Zero-Shot 3D Scene Understanding
浏览论文内容
中文总结 AI 辅助
针对现有开放词汇零样本3D场景理解模型难以处理流式RGB-D输入、易受2D分割掩码噪声影响的问题,提出无训练框架Stream3Dv2,通过几何-语义融合等策略提升性能,在相关任务上优于基准,可与LLM智能体集成用于开放世界具身智能。
中文摘要 AI 辅助
近年来,基于视觉基础模型的开放词汇零样本3D场景理解已成为数据密集型监督方法的有前景替代方案。然而,将这些模型部署到真实场景中受到严重阻碍,原因在于它们无法高效处理流式RGB-D输入,且固有地易受噪声2D分割掩码的影响。为解决这些关键局限,我们提出Stream3Dv2,这是一种专为鲁棒流式3D感知设计的新型无训练框架。Stream3Dv2通过原创的嵌套局部到历史架构处理序列数据,捕获多视图一致性同时规避高计算开销,以支持及时响应。其核心是我们引入的综合几何-语义融合机制,该机制通过显式利用语义指导并将3D分割表述为点集合并与划分问题,解决几何噪声和语义歧义。此外,我们提出了一种创新的基于流形距离的点云优化策略,该方法利用局部流形图进行点到流形优化,缓解欧氏距离度量导致的边界划分失败,并采用几何边界框动态激活和更新历史实例,以实现快速的流形到流形优化。在公共数据集上的大量实验表明,Stream3Dv2在基础开放词汇流式3D分割和检测任务中始终优于现有基准。最后,我们展示将我们的框架与基于大语言模型(LLM)的智能体集成可实现高级语言驱动的3D场景理解,凸显其在开放世界具身智能领域的潜力。代码将在此httpsURL更新。
英文摘要
Recently, open-vocabulary zero-shot 3D scene understanding using vision foundation models has emerged as a promising alternative to data-intensive supervised methods. However, deploying these models in real-world scenarios is severely hindered by their inability to efficiently handle streaming RGB-D inputs and their inherent vulnerability to noise 2D segmentation masks. To address these critical limitations, we propose Stream3Dv2, a novel training-free framework designed for robust streaming 3D perception. Stream3Dv2 processes sequential data through an original nested local-to-historical architecture, capturing multi-view consistency while circumventing the high computational overhead so as to support timely responses. At its core, we introduce a comprehensive geometric-semantic fusion mechanism that resolves geometric noise and semantic ambiguity by explicitly utilizing semantic guidance and formulating 3D segmentation as solving point-and-set merging and partitioning problems. Furthermore, we present an innovative manifold-distance-based point cloud refinement strategy. This approach leverages local manifold graphs for point-to-manifold optimization that mitigates the boundary delineation failures caused by Euclidean-distance metrics, and employs geometric bounding boxes to dynamically activate and update historical instances for achieving rapid manifold-to-manifold refinement. Extensive experiments on public datasets demonstrate that Stream3Dv2 consistently outperforms existing baselines in foundational open-vocabulary streaming 3D segmentation and detection. Finally, we show that integrating our framework with an LLM-based agent enables advanced language-driven 3D scene understanding, underscoring its potential for open-world embodied intelligence. Code will be updated at https://github.com/SubmissionsIn/Stream3D.
发表机构
- Singapore University of Technology and Design(新加坡科技设计大学)
机构由 AI 辅助整理,请以论文原文为准。