发表机构
Horizon Robotics; WuwenAI; Southeast University; Huazhong University of Science and Technology(地平线机器人; 悟文智能; 东南大学; 华中科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有操作基准无法充分检验机器人指令跟随能力的问题,提出文本不可或缺的InstructMove基准,将指令跟随分解为多环节,实验表明其可诊断视觉捷径且模拟数据能提升现实操作性能。
AI 中文摘要
视觉-语言-动作(VLA)模型通过将机器人动作建立在自然语言指令的条件下,使通用机器人操作日益可行。此类通用性的关键检验在于策略是否真正遵循语言指令。然而,许多操作基准未能充分检验这一能力:目标对象或目的地往往在视觉上显著或具备唯一可行性,使得策略无需理解指令即可成功。我们认为,指令跟随评估应是文本不可或缺的:多个动作在视觉和物理上均合理,而仅一个动作符合语言指令。我们提出InstructMove,一种用于指令跟随操作的文本不可或缺基准。InstructMove在带有语义干扰项的抓取-放置场景中实现了这一原则,将指令跟随分解为类别识别、属性判别、空间推理和组合式抓取-放置。InstructMove支持使用InstructMove训练数据和保留的评估任务的训练-评估协议,还提供语言依赖性诊断功能。对代表性VLA策略的实验表明,InstructMove提供了一个受控测试平台,用于诊断视觉捷径,且InstructMove模拟数据可提升现实世界指令跟随操作的性能。代码:this https URL
英文摘要
Vision-language-action (VLA) models have made general-purpose robot manipulation increasingly plausible by conditioning robot actions on natural-language instructions. A key test of such generality is whether policies actually follow language instructions. Yet many manipulation benchmarks leave this ability underdetermined: the intended object or destination is often visually salient or uniquely feasible, allowing policies to succeed without grounding the instruction. We argue that instruction-following evaluation should be text-indispensable: multiple actions should be visually and physically plausible, while only one should be consistent with the language instruction. We introduce InstructMove, a text-indispensable benchmark for instruction-following manipulation. InstructMove instantiates this principle in pick-and-place scenes with semantic distractors, decomposing instruction following into category identification, attribute discrimination, spatial reasoning, and compositional pick-and-place. InstructMove supports a train-eval protocol with InstructMove training data and held-out evaluation tasks, with additional diagnostics for language dependence. Experiments with representative VLA policies show that InstructMove provides a controlled testbed for diagnosing visual shortcuts and that InstructMove simulation data can improve real-world instruction-following manipulation performance. Code: https://github.com/HorizonRobotics/RoboOrchardSim
Comments22 pages