arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RefGuard:基于联合目标-锚点-坐标系定位的身份感知语言引导机器人操作

RefGuard: Identity-Aware Language-Guided Robot Manipulation via Joint Target-Anchor-Frame Grounding

Lan Wei, Kangyi Lu, Yongchen Wang, Chenmeng Bi, Qi Chen, Hanlin Niu, Yip Fun Yeung, Dandan Zhang

arXiv 2609.06221首次发表:更新:

发表机构

Imperial College London; RACE, United Kingdom Atomic Energy Authority; Oxford Robotics Institute, University of Oxford; Avant Intelligence Ltd(帝国理工学院; 英国原子能管理局RACE; 牛津大学牛津机器人研究所; Avant Intelligence有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

RefGuard提出身份感知定位框架,通过联合维护目标、锚点和坐标系的概率后验并延迟决策,在杂乱场景中消除语言引导操作的身份切换,显著提升执行准确率。

AI 中文摘要

视觉-语言-动作(VLA)模型已大幅推进语言引导的机器人操作,但可靠执行仍取决于识别指令所指的物理对象。在包含重复对象、模糊锚点或依赖坐标系的方位词的杂乱场景中,机器人可能会在语义兼容但非预期的实例上执行几何上有效的动作;我们将这种失败称为身份切换。指代对象由三个耦合的潜在变量共同决定:目标、锚点和参考坐标系,因此在执行前对其中任何一个做出承诺,都会将残余歧义转化为无声且不可逆的错误。我们提出RefGuard,一个身份感知的定位框架,通过维护所有三个变量的联合后验来延迟承诺。RefGuard从RGB-D观测构建基于坐标系的对象中心场景图,将坐标系无关的几何与方向关系分离,并将后验路由至决策策略,该策略可执行、澄清、重新观察或中止。在真实的UF850机械臂上,RefGuard在所有歧义压力试验中均未记录身份切换,并在其中90.0%的试验中正确执行,而微调的VLA和基于LLM(大语言模型)的基线在同一试验中有33-46%发生身份切换,同时在无歧义场景中保持93.3%的成功率,并在86.7%的试验中从定位后场景变化中恢复。在包含3200个回合的程序化套件中,相较于在目标前承诺锚点和坐标系的消融方法,RefGuard将可解指令的正确执行率从56.6%提升至80.5%,同时更少地推迟(19.5%对43.4%)。

英文摘要

Vision-language-action (VLA) models have substantially advanced language-guided robot manipulation, yet reliable execution still hinges on identifying which physical object an instruction refers to. In cluttered scenes containing repeated objects, ambiguous anchors, or frame-dependent spatial terms, a robot can execute a geometrically valid action on a semantically compatible but unintended instance; we call this failure an identity switch. The referent is jointly determined by three coupled latent variables: the target, the anchor, and the reference frame, so committing to any one of them before execution turns residual ambiguity into a silent and irreversible error. We propose RefGuard, an identity-aware grounding framework that delays commitment by maintaining a joint posterior over all three variables. RefGuard builds a frame-conditioned object-centric scene graph from RGB-D observations, separating frame-independent geometry from directional relations, and routes the posterior through a decision policy that executes, clarifies, reobserves, or aborts. On a real UF850 arm, RefGuard records no identity switch on any ambiguity-stress trial and executes correctly on 90.0% of them, whereas fine-tuned VLA and LLM (Large Language Model)-based baselines switch identity in 33-46% of the same trials, while retaining 93.3% success on unambiguous scenes and recovering from post-grounding scene changes in 86.7% of trials. On a 3200-episode procedural suite, it raises correct execution on solvable instructions from 56.6% to 80.5% over the ablation that commits to the anchor and frame before the target, while deferring less often (19.5% vs. 43.4%).

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑