arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RoboFollow:揭示具身智能体中的指令跟随幻象

RoboFollow: Unveiling the Instruction Following Mirage in Embodied Agents

Chang Guo, Yukun Xie, Bohan Tan, Zheng Chang, Zhaokai Yin, Qianli Ma, Yingqiao Wang, Chao Liang, Zhipeng Zhang

arXiv 2609.25636首次发表:更新:

发表机构

AutoLab, School of Artificial Intelligence, Shanghai Jiao Tong University; Research Lab, Anyverse Dynamics(上海交通大学人工智能学院AutoLab; Anyverse Dynamics研究实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

RoboFollow通过高场景熵诊断基准揭示具身智能体指令跟随的幻象,证明低场景熵下高成功率掩盖了语言依赖不足,且现有缓解措施无法弥补该瓶颈。

AI 中文摘要

现代具身智能体取得了令人瞩目的成功率,但其实际的指令跟随能力远比这些数字所显示的要弱。我们将这一幻象追溯到一个我们称之为低场景熵的结构性属性:当视觉场景仅允许一个有效任务时,语言变得冗余,策略可以在几乎不使用语言的情况下获得高分。我们引入了RoboFollow,一个具有三项原则的诊断基准:(1)高场景熵:每个训练场景支持多个运动学上不同的任务分支,使得仅凭视觉不足以完成任务,从而迫使策略依赖语言。(2)分层诊断协议:一个四级协议(L0--L3)逐步扰动视觉布局和语义,探测等效指令是否产生一致的行为,以及不同指令是否在空间关系、属性、轨迹约束和逻辑方面产生可区分的行为。(3)混淆控制诊断:我们简化交互对象,将动作限制在训练过的动作库内,并报告分阶段的意图和执行分数,从而将理解与运动执行分离。对九个VLA和WAM策略的评估表明,在我们的微调设置下,强大的L0性能(若达到)并不能可靠地迁移到L1--L3。代表性的缓解措施,包括更强的VLM骨干网络、QA联合训练、LangForce和Classifier-Free Guidance,均未能缩小这一差距。RoboFollow揭示了真正的指令跟随是一个关键且被忽视的瓶颈。代码和数据集可在以下网址获取:此https URL和此https URL。

英文摘要

Modern embodied agents achieve impressive success rates, yet their actual instruction-following ability is far weaker than these numbers suggest. We trace this illusion to a structural property we term low scene entropy: when a visual scene admits only one valid task, language becomes redundant and a policy can score highly while barely using it. We introduce RoboFollow, a diagnostic benchmark with three principles: (1) High Scene Entropy: each training scene supports multiple kinematically distinct task branches, making vision alone insufficient and forcing reliance on language. (2) Hierarchical Diagnostic Protocol: a four-level protocol (L0--L3) progressively perturbs visual layout and semantics, probing whether equivalent instructions yield consistent behavior and distinct ones yield discriminable behavior across spatial relations, attributes, trajectory constraints, and logic. (3) Confound-Controlled Diagnosis: we simplify interaction objects, restrict actions to the trained repertoire and report stage-wise Intent and Execution scores, isolating comprehension from motor execution. Evaluation of nine VLA and WAM policies shows that strong L0 performance, where attained, does not reliably transfer to L1--L3 under our fine-tuning setup. Representative mitigations, including stronger VLM backbones, QA co-training, LangForce, and Classifier-Free Guidance, all fail to close this gap. RoboFollow exposes genuine instruction following as a critical, overlooked bottleneck. Code and dataset are available at https://github.com/AutoLab-SAI-SJTU/RoboFollow and https://huggingface.co/datasets/AutoLab-SJTU/robofollow-data.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑