arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

InstructMove:一种指令跟随操作的文本不可或缺基准

InstructMove: A Text-Indispensable Benchmark for Instruction-Following Manipulation

Mengao Zhao, Ziang Li, Chaodong Huang, Mengchen Ma, Haoyi Jiang, Yiwei Jin, Xinjie Wang, Yun Du, Xuewu Lin, Taojun Ding, Hongyu Xie, Jackson Jiang, Chunlei Yu, Kaihua Zhang, Lichao Huang, Liu Liu, Tianwei Lin, Zhizhong Su

arXiv 2608.22990首次发表:更新:

发表机构

Horizon Robotics; WuwenAI; Southeast University; Huazhong University of Science and Technology(地平线机器人; 悟文智能; 东南大学; 华中科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有操作基准无法充分检验机器人指令跟随能力的问题,提出文本不可或缺的InstructMove基准,将指令跟随分解为多环节,实验表明其可诊断视觉捷径且模拟数据能提升现实操作性能。

AI 中文摘要

视觉-语言-动作(VLA)模型通过将机器人动作建立在自然语言指令的条件下,使通用机器人操作日益可行。此类通用性的关键检验在于策略是否真正遵循语言指令。然而,许多操作基准未能充分检验这一能力:目标对象或目的地往往在视觉上显著或具备唯一可行性,使得策略无需理解指令即可成功。我们认为,指令跟随评估应是文本不可或缺的:多个动作在视觉和物理上均合理,而仅一个动作符合语言指令。我们提出InstructMove,一种用于指令跟随操作的文本不可或缺基准。InstructMove在带有语义干扰项的抓取-放置场景中实现了这一原则,将指令跟随分解为类别识别、属性判别、空间推理和组合式抓取-放置。InstructMove支持使用InstructMove训练数据和保留的评估任务的训练-评估协议,还提供语言依赖性诊断功能。对代表性VLA策略的实验表明,InstructMove提供了一个受控测试平台,用于诊断视觉捷径,且InstructMove模拟数据可提升现实世界指令跟随操作的性能。代码:this https URL

英文摘要

Vision-language-action (VLA) models have made general-purpose robot manipulation increasingly plausible by conditioning robot actions on natural-language instructions. A key test of such generality is whether policies actually follow language instructions. Yet many manipulation benchmarks leave this ability underdetermined: the intended object or destination is often visually salient or uniquely feasible, allowing policies to succeed without grounding the instruction. We argue that instruction-following evaluation should be text-indispensable: multiple actions should be visually and physically plausible, while only one should be consistent with the language instruction. We introduce InstructMove, a text-indispensable benchmark for instruction-following manipulation. InstructMove instantiates this principle in pick-and-place scenes with semantic distractors, decomposing instruction following into category identification, attribute discrimination, spatial reasoning, and compositional pick-and-place. InstructMove supports a train-eval protocol with InstructMove training data and held-out evaluation tasks, with additional diagnostics for language dependence. Experiments with representative VLA policies show that InstructMove provides a controlled testbed for diagnosing visual shortcuts and that InstructMove simulation data can improve real-world instruction-following manipulation performance. Code: https://github.com/HorizonRobotics/RoboOrchardSim

Comments22 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑