发表机构
University of Cambridge; Massachusetts Institute of Technology; Hong Kong University of Science and Technology; University of Oxford; Harvard University(剑桥大学; 麻省理工学院; 香港科技大学; 牛津大学; 哈佛大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出一种受大脑启发的零样本分层框架,通过显式状态推理、原子动作组合、成本排序和闭环验证,在机器人任务推理与执行中超越现有基准。
AI 中文摘要
遵循开放式语言指令的机器人需要将语义意图与视觉场景理解、几何可行性、物体状态和物理交互条件联系起来。端到端的视觉-语言-动作策略提高了跨任务泛化能力,但它们通常直接将视觉和语言输入映射到机器人动作,缺乏用于长时程任务分解、物理验证和恢复的显式结构。我们提出\method,一种零样本分层框架,其功能上受人类大脑角色分工的启发,包括视觉感知与状态推断、基于共享原子动作库的语言接地与动作序列生成、基于成本的方案选择,以及真实机器人执行与验证。该框架将指令接地到显式物体状态,将可复用的原子动作组合成任务条件序列,按执行成本对备选序列排序,并根据更新的观察验证中间物理结果。在评估中,\method在平坦和不规则初始布局条件下均完成了10/10次干净棋盘试验、10/10次抓取放置试验和4/5次金字塔堆叠试验;相应的平均任务进度分别为$99.03\\%$、$100.00\\%$和$96.67\\%$。在所有评估条件下,\method的成功率均高于ReKep、Dream2Flow和$\pi_{0.5}$基准,证明了结合显式物体状态推理、组合原子动作、基于成本的方案选择和闭环执行验证的有效性。
英文摘要
Robots that follow open-ended language instructions need to connect semantic intent to visual scene understanding, geometric feasibility, object states, and physical interaction conditions. End-to-end Vision-Language-Action policies have improved cross-task generalization, but they typically map visual and language inputs directly to robot actions, leaving limited explicit structure for long-horizon decomposition, physical verification, and recovery. We present \method, a zero-shot hierarchical framework functionally inspired by the division of roles in the human brain, comprising visual perception and state inference, language grounding and action-sequence generation from a shared atomic action library, cost-based plan selection, and real-robot execution and verification. The framework grounds commands in explicit object states, composes reusable atomic actions into task-conditioned sequences, ranks alternative sequences by execution cost, and verifies intermediate physical outcomes from refreshed observations. In the evaluation, \method{} completes 10/10 clean board trials, 10/10 pick-and-place trials, and 4/5 pyramid stacking trials for both the flat and irregular initial-layout conditions; the corresponding mean task progress is $99.03\%$, $100.00\%$, and $96.67\%$ respectively. Across all evaluated conditions, \method{} achieves higher success rates than ReKep, Dream2Flow, and $π_{0.5}$ benchmarks, demonstrating the effectiveness of combining explicit object-state reasoning, compositional atomic actions, cost-based plan selection, and closed-loop execution verification.
Comments10 pages, 5 figures