发表机构
National University of Singapore; University of California, Los Angeles(新加坡国立大学; 加州大学洛杉矶分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
MIGU通过融合几何与语义不确定性,实现多模态指令接地,提升操作规划性能,在真实基准上优于基线。
AI 中文摘要
理解自然的人类指令对于在人类中心环境中部署机器人至关重要。我们研究多模态指令接地,其中语言和手势提供互补但不确定的线索。我们提出MIGU,一个模块化框架,将语义和几何证据结合成统一的接地信念,并将其连接到操作规划。MIGU通过眼-手指几何传播视线方向和深度不确定性,同时考虑手方向估计误差,构建3D几何似然。视觉语言模型(VLM)提供候选对象和区域的语义先验,通过贝叶斯启发式融合与几何似然结合。由此产生的信念支持行为规划,要么直接进行下游规划,要么请求澄清。接地的目标随后定义移动操作和桌面任务与运动规划的目标。在真实世界基准上,MIGU优于所有评估的基线,而消融研究支持显式多模态不确定性建模的益处。项目网站:此HTTP URL
英文摘要
Understanding natural human instructions is crucial for deploying robots in human-centric environments. We study multimodal instruction grounding, where language and gesture provide complementary but uncertain cues. We present MIGU, a modular framework that combines semantic and geometric evidence into a unified grounding belief and connects it to manipulation planning. MIGU constructs a 3D geometric likelihood by propagating viewing-direction and depth uncertainty through eye-finger geometry while accounting for hand-direction estimation error. A vision-language model (VLM) provides semantic priors over candidate objects and regions, which are combined with the geometric likelihood through Bayes-inspired fusion. The resulting belief supports behavior planning to either proceed directly to downstream planning or request clarification. Grounded targets then define goals for mobile manipulation and tabletop task-and-motion planning. On a real-world benchmark, MIGU outperforms all evaluated baselines, while ablations support the benefit of explicit multimodal uncertainty modeling. Project website: multimodal-instruction.github.io