arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MonoVoc:解耦几何与语义的轻量单目开放词汇3D高斯模型

MonoVoc: Decoupling Geometry and Semantics for Lightweight Monocular Open-Vocabulary 3D Gaussians

Pouya Ardekhani, Zahra Dehghanian, Morteza Abolghasemi, Hamid R. Rabiee

arXiv 2607.28300首次发表:更新:

发表机构

Sharif University of Technology(谢里夫理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

MonoVoc通过解耦几何与语义的无训练流水线,以轻量模块化后处理框架实现单目开放词汇3D高斯建模,内存较SOTA降一个数量级,适配日常单目视频的3D场景理解任务。

AI 中文摘要

开放词汇3D场景理解是下一代交互系统的核心,支持用户通过自然语言直观查询、导航重建环境。但现有3D高斯框架受限于严格的多视图采集要求、高昂的场景特定优化成本,以及存储密集语言特征的巨大内存开销。我们提出一种新颖的无训练流水线,通过显式解耦3D几何重建与语义集成,彻底重构该范式:输入标准单目视频序列,即可高效输出紧凑、高可解释、可全搜索的对象级语义高斯地图。我们未将沉重的语言嵌入与建图循环绑定,而是独立提取几何,通过轻量模块化后处理框架锚定语义。在Replica数据集上的大量评估表明,这种解耦架构保留了强渲染保真度与具竞争力的分割精度;关键的是,相比SOTA基线,我们的方法用模块化对象级语义嵌入替换密集的逐高斯存储,内存用量降低一个数量级,为日常单目视频的开放词汇3D检索与问答提供了高效、可扩展且实用的解决方案。

英文摘要

Open vocabulary 3D scene understanding is essential for next-generation interactive systems, empowering users to intuitively query and navigate reconstructed environments using natural language. However, current 3D Gaussian frameworks are often bottlenecked by restrictive multiview capture requirements, costly scene-specific optimization, and the massive memory overhead of storing dense language features. We present a novel, training-free pipeline that fundamentally reimagines this paradigm by explicitly decoupling 3D geometric reconstruction from semantic integration. Given a standard monocular video sequence as input, our method efficiently outputs a compact, highly interpretable, and fully searchable object-level semantic Gaussian map. Rather than entangling heavy language embeddings within the mapping loop, we extract geometry independently and ground semantics through a lightweight, modular post-processing framework. Extensive evaluations on the Replica dataset demonstrate that this decoupled architecture preserves strong rendering fidelity and competitive segmentation accuracy. Crucially, by replacing dense per-Gaussian storage with modular, object-level semantic embeddings, our approach delivers an order-of-magnitude reduction in memory usage compared to SOTA baselines. This provides a highly efficient, scalable, and practical solution for open-vocabulary 3D retrieval and question answering directly from everyday monocular video.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑