基于语义三维高斯溅射的开放词汇移动操作具身多模态定位
Embodied Multimodal Grounding for Open-Vocabulary Mobile Manipulation via Semantic 3D Gaussian Splatting
浏览论文内容
中文总结 AI 辅助
本文针对家庭场景开放词汇移动操作的目标定位问题,提出整合Semantic-3DGS的具身多模态框架,经50次真实机器人试验,在长 horizon、杂乱场景等任务中优于PointVLA等基线方法,提升了操作鲁棒性。
中文摘要 AI 辅助
具身移动操作要求在执行前对齐语言、视觉观测、三维场景结构与动作可行性。本文研究本地家庭工作空间中少样本操作的开放词汇目标定位,提出一种具身多模态定位框架,整合主动多视图语义三维高斯溅射(Semantic-3DGS)、可达性感知的基座定位以及基于扩散的视觉-语言-动作策略。任务驱动的局部Semantic-3DGS作为共享接口,连接主动感知、语言条件三维定位、障碍物感知场景推理、基座准备及动作模型的语义条件化。为保留预训练动作先验,仅将三维语义线索注入后期动作专家模块。在针对代表性视觉-语言-动作(VLA)方法的扩展50次真实机器人试验评估中,完整系统实现60%的长 horizon 成功率,而PointVLA为40%、DexVLA为28%;在重度杂乱操作中达到74%的成功率,单视图变体为52%、PointVLA为46%;在75cm高度偏移下仍保持75%的成功率,且消除了光诱导的错误抓取。这些结果表明,显式、可刷新的三维语义定位可提升在杂乱、遮挡、视角变化及具身约束下的鲁棒性。
英文摘要
Embodied mobile manipulation requires language, visual observations, three-dimensional scene structure, and action feasibility to be aligned before execution. We study open-vocabulary target grounding with few-shot manipulation in local household workspaces and present an embodied multimodal grounding framework that integrates active multi-view Semantic 3D Gaussian Splatting (Semantic-3DGS), reachability-aware base positioning, and a diffusion-based vision-language-action policy. A task-driven local Semantic-3DGS serves as a shared interface across active sensing, language-conditioned 3D localization, obstacle-aware scene reasoning, base preparation, and semantic conditioning of the action model. To preserve pretrained action priors, the 3D semantic cues are injected only into the late action-expert blocks. In expanded 50-trial real-robot evaluations against representative vision-language-action (VLA) approaches, the full system achieves 60% long-horizon success compared with 40% for PointVLA and 28% for DexVLA, and reaches 74% success in heavily cluttered manipulation compared with 52% for the single-view variant and 46% for PointVLA. It also maintains 75% success under a 75 cm height shift and eliminates photo-induced false grasps. These results indicate that explicit, refreshable 3D semantic grounding can improve robustness under clutter, occlusion, viewpoint variation, and embodiment constraints.
发表机构
- The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
- Midea Group(美的集团)
- The Hong Kong University of Science and Technology(香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。