LEGO:分层语言高斯溅射
LEGO: Leveled Language Gaussian Splatting
浏览论文内容
中文总结 AI 辅助
LEGO是一种新方法,通过将多视角SAM粒度转为3D一致层级,结合CLIP嵌入构建分层语言场景图,在3D分割基准上取得最优性能,可赋能大语言模型的空间推理与视觉定位。
中文摘要 AI 辅助
我们推出LEGO以实现高级开放词汇场景理解。除基础概念识别外,其核心创新在于捕捉场景内的内在语义层级,例如“花盆→花束→花蕾→花瓣”的谱系。尽管SAM等基础模型可识别2D中的多粒度结构,但其划分严格受视角限制,缺乏跨视角一致性。LEGO自适应地将不稳定的多视角SAM粒度重新分级为统一的3D一致层级,为3D场景的结构连贯多级别分割提供精确监督。通过用CLIP嵌入锚定这些分割,LEGO恢复了跨层级的开放词汇语义逻辑。此外,通过整合空间关系,我们将这些分割提升为分层语言场景图,有效赋能大语言模型执行复杂、上下文感知的空间推理和精确视觉定位。实验结果表明,LEGO在可提示和开放词汇3D分割基准上均达到新的最优性能,展现出先进的分层场景分解和上下文感知空间推理能力。
英文摘要
We introduce LEGO for advanced open-vocabulary scene understanding. Beyond basic concept recognition, its core innovation lies in capturing the intrinsic semantic hierarchies within the scene, such as the "flowerpot -> bouquet -> bud -> petal" lineage. While foundation models like SAM can identify multi-granular structures in 2D, their partitions are strictly perspective-bound and lack cross-view consensus. LEGO self-adaptively re-grades volatile multi-view SAM granularities into a unified, 3D-consistent hierarchy. This provides precise supervision for the structurally coherent, multi-level segmentation of 3D scenes. By grounding these segments with CLIP embeddings, LEGO recovers open-vocabulary semantic logic across hierarchical levels. Furthermore, by incorporating spatial relationships, we elevate these segments into level-wise language scene graphs, effectively empowering Large Language Models to perform complex, context-aware spatial reasoning and precise visual grounding. Experimental results demonstrate that LEGO establishes new state-of-the-art performance across both promptable and open-vocabulary 3D segmentation benchmarks, exhibiting advanced hierarchical scene decomposition and context-aware spatial reasoning.
发表机构
- Wuhan University(武汉大学)
- Hong Kong University of Science and Technology(香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。