HumanCLAW:视觉-语言模型能否通过实体身体执行动作
HumanCLAW: Can Vision-Language Models Act Through a Body?
- Meta
- Nanyang Technological University(南洋理工大学)
- University of Washington(华盛顿大学)
- Brown University(布朗大学)
- Northwestern University(西北大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出HumanCLAW框架与HumanCLAW-Bench基准,测试9个先进VLM发现其均未解决具身动作任务,最优仅16.8%成功率,核心缺失具身自我意识。
AI中文摘要:
评估视觉-语言模型(VLM)能否通过物理实体身体执行动作颇具挑战,因为动作结果会耦合VLM的决策与运动控制。当任务失败时,很难区分是VLM做出了错误选择,还是运动控制器执行失败,例如失去平衡摔倒。本研究提出HumanCLAW,这是一个将动作决策与底层执行解耦的评估框架。每一步,搭载的现成VLM会发出原子技能指令,该指令被转换为亚秒级的连续全身运动片段,伴随重力、碰撞等真实物理后果。如此,实体身体可在物理世界自由动作,而执行侧的干扰、平衡与运动误差被排除,可测量的仅为模型的动作智能:其逐时刻选择身体下一步应执行动作的能力。基于该框架,构建了HumanCLAW-Bench:在41个室内场景中包含1218个长视野、第一人称视角的查找-导航-交互任务 episode。测试9个最先进的VLM后发现,没有一个能解决该基准任务,最优模型仅达到16.8%的成功率。研究还指出,识别目标并非瓶颈,当前VLM缺失的是具身自我意识:它们会丢失自身身体的状态,无法判断身体位置、是否到达目标或是否碰到障碍物。
英文摘要:
Evaluating whether a vision-language model (VLM) can act through a physical body is challenging. The outcome of an action couples the VLM's decision with motor control. When a task fails, it is hard to tell whether the VLM made a bad choice or the motor controller simply failed to execute it, e.g., losing balance and falling. In this work, we introduce HumanCLAW, an evaluation framework that decouples action decision-making from low-level execution. At every step, a harnessed, off-the-shelf VLM issues an atomic skill command, and the command is translated into a sub-second chunk of continuous full-body motion with real physical consequences, including gravity and collisions. The body can therefore act freely in the physical world, while execution-side disturbances, balance and motor errors, are factored out. What remains measurable is the model's action intelligence: its moment-to-moment choice of what the body should execute next. Based on this framework, we build HumanCLAW-Bench: 1,218 long-horizon, egocentric find-navigate-interact episodes across 41 indoor scenes. We test nine state-of-the-art VLMs and find that none solves the benchmark; the best model reaches only a 16.8% success rate. Recognizing the target is not the bottleneck. What current VLMs lack is embodied self-awareness: they lose track of their own body, failing to tell where it is, whether it has reached the goal, or whether it has hit an obstacle.