arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34384cs.ROcs.CV

RoboIRGBench:视觉-语言-动作模型中的隐式指代接地基准

RoboIRGBench: Benchmarking Implicit Referential Grounding in Vision-Language-Action Models

Aernaer Akelijiang, Jiannan Li, Zhineng Chen, Jingjing Chen, Bin Zhu

首次发表
浏览论文内容

中文总结 AI 辅助

针对现有基准假设指令显式指定任务信息的问题,提出RoboIRG-Bench基准,系统评估VLA模型在隐式指代接地上的能力,揭示显式与隐式指令间的性能差距,并强调需提升模型整合语言、感知、推理与动作的鲁棒性。

中文摘要 AI 辅助

视觉-语言-动作(VLA)模型在机器人操作中展现出强大的能力,然而现有基准通常假设任务相关信息在指令中被明确指定。但在实践中,人类经常隐式地指代对象、数量和关系,要求机器人从语言和感知上下文中恢复预期目标。我们将这种能力研究为隐式指代接地(IRG),并引入RoboIRG-Bench,一个旨在系统评估该能力的操作基准。基于RoboMME构建,RoboIRG-Bench包含源自11个任务的40个变体,涵盖四个挑战,包括直接、推理中介、空间和上下文指代接地。由于IRG通常需要保留和检索先前建立的上下文,我们评估了跨越不同记忆机制的代表性VLA模型。我们的评估揭示了显著的指代鲁棒性差距。在显式指令下表现良好的模型,当相同的任务相关信息必须从上下文中恢复时,性能会急剧下降。推理中介和空间指代尤其具有挑战性,而使用外部VLM的模型表现出更强的鲁棒性但仍存在显著失败。此外,用更强的模型替换外部VLM并不能消除这些差距。我们进一步在Franka Research 3机械臂上验证了这些发现,在真实世界操作中差距持续存在,表现为错误的指代接地和下游执行失败。这些结果确立了IRG作为可靠机器人指令跟随中一个独特且未被充分探索的能力,并强调需要能够稳健整合语言、感知、推理和动作的VLA模型。

英文摘要

Vision-Language-Action (VLA) models have shown strong capabilities in robotic manipulation, yet existing benchmarks typically assume that task-relevant information is explicitly specified in the instruction. In practice, however, humans frequently refer to objects, quantities, and relations implicitly, requiring robots to recover the intended target from linguistic and perceptual context. We study this capability as Implicit Referential Grounding (IRG) and introduce RoboIRG-Bench, a manipulation benchmark designed to systematically evaluate it. Built upon RoboMME, RoboIRG-Bench contains 40 variants derived from 11 tasks and covers four challenges, including direct, reasoning-mediated, spatial, and contextual referential grounding. As IRG often requires retaining and retrieving previously established context, we evaluate representative VLAs spanning different memory mechanisms. Our evaluation reveals a noticeable referential robustness gap. Models that perform well under explicit instructions can degrade sharply when the same task-relevant information must be recovered from context. Reasoning-mediated and spatial references are particularly challenging, while models using external VLMs show greater robustness but still exhibit significant failures. Moreover, replacing the external VLM with a stronger model does not eliminate these gaps. We further validate these findings on a Franka Research 3 robot arm, where the gap persists under real-world manipulation and manifests as both incorrect referent grounding and downstream execution failures. These results establish IRG as a distinct and underexplored capability for reliable robotic instruction following and highlight the need for VLAs that can robustly integrate language, perception, reasoning, and action.

发表机构

  • Singapore Management University(新加坡管理大学)
  • Fudan University(复旦大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑