AI 中文总结
IMBench旨在评估直观机器人操作这一综合能力,其任务含物理结构推断与动作序列生成。通过35个任务等进行实验,发现视觉语言模型和视觉 - 语言 - 动作模型的不足,为评估和发展物理智能提供基准。
AI 中文摘要
人类结合推理和运动控制来解决各种约束下的复杂操作任务,这种能力被称为直观操作。现有基准测试未能捕捉到这种整合。IMBench旨在评估直观操作这一跨越感知、物理推理、动作生成和迭代执行的综合能力。其任务要求模型推断与任务相关的物理结构并在明确约束下生成可行动作序列。实验揭示了视觉语言模型和视觉 - 语言 - 动作模型的不足,表明直观操作是当前基础模型和通用机器人策略中缺失的维度,IMBench朝着评估和实现更综合、自适应的物理智能迈进。
英文摘要
Humans combine reasoning and motor control to solve complex manipulation tasks under diverse constraints. They build an understanding of the physical world that helps them convert reasoning into actions and quickly adapt to new scenes, tasks, and rules. We refer to this capability as intuitive manipulation. Existing benchmarks fail to capture this integration: they evaluate physical reasoning in isolation from execution, or measure policy performance without requiring explicit reasoning. We introduce IMBENCH, a benchmark designed to evaluate intuitive manipulation as an integrated capability spanning perception, physical reasoning, action generation, and iterative execution. Our tasks require models to infer task-relevant physical structure and generate feasible action sequences under explicit constraints, including contact-rich manipulation, tool use, and multi-stage dependencies. We introduce a benchmark of 35 tasks, 14K filtered trajectories, and scalable tools for generating diverse scenarios. Experiments reveal a consistent gap: vision language models show partial physical reasoning ability but fail to produce executable plans, while state-of-the-art vision-language-action models struggle to satisfy task constraints and generalize across scenarios. These results identify intuitive manipulation as a missing axis in current foundation models and generalist robot policies, and position IMBENCH as a step toward evaluating and enabling more integrated, adaptive physical intelligence.
CommentsAccepted to SemRob Workshop, RSS 2026. Project Website: https://imbench.org/