arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RoboAware:基于反事实结果学习具身技能的协调方法

RoboAware: Learning to Coordinate Embodied Skills from Counterfactual Outcomes

Bohan Zhou, Xingbei Chen, Emily Huang, Weilin Ruan, Haojian Huang, Yehang Zhang, Zexi Li, Wenqian Li, Qize Yu, Zetian Song, Leyi Wu, Jinghao Li, Mingxuan Song, Xinrun Xu, Zongyang Qiu, Yangkai Wei, Tianyi Zhang, Kaiwen Zhou, Yinchuan Li, James Cheng

arXiv 2610.11480首次发表:更新:

AI 中文总结

RoboAware是一种基于反事实结果学习的具身技能协调方法,通过$P^5$架构、SCB和EAL技术,在100个任务上总体成功率达77.0%,优于现有基线方法。

AI 中文摘要

具身编码智能体可将模块化机器人技能与冻结的端到端策略相结合,但有效的组合需要预判当前物理状态下哪种策略族会成功。我们提出RoboAware,它在编码智能体的技能编排基础上,仅从反事实结果中学习状态条件下的责任协调器。受REPL成功的启发,我们提出$P^5$架构并基于其构建分层马尔可夫决策过程(MDP)。$P^5$将技能统一组织为五个语义阶段,定义了可比较责任的位置。为解决现有工作中缺乏反事实分支结果的问题,我们引入状态锁定反事实分支(SCB),该方法恢复相同的训练状态,以从每个可接受的策略族生成并执行代码块,从而揭示选定分支经验未观察到的结果。在此基础上,我们提出执行感知学习(EAL),它将蒙特卡洛树搜索与Q学习相结合,以将这些结果提炼为策略族条件值。部署时,协调器根据可观察上下文选择策略族,冻结的编码智能体生成下一个本地代码块。对100个任务的全面单回合评估显示,RoboAware的总体成功率达77.0%,在RoboSuite上达到90.0%的SOTA平均值、在多样化LIBERO-Pro任务集群上达到73.8%、在具有挑战性的RoboTwin双臂任务上达到90.0%,优于现有的代码即策略和VLA-harness基线方法。

英文摘要

Embodied coding agents can combine modular robot skills with frozen end-to-end policies, yet effective composition requires anticipating which policy family will succeed in the current physical state. We present RoboAware, which builds on coding agents' skill orchestration by learning only a state-conditioned responsibility coordinator from counterfactual outcomes. Inspired by the success of REPL, we propose the $P^5$ schema and formulate a hierarchical MDP based on it. $P^5$ organizes skills uniformly into five semantic stages, defining where responsibility can be compared. To address the lack of counterfactual branch outcomes in existing work, we introduce State-Locked Counterfactual Branching (SCB), which restores the same training state to generate and execute a code block from each admissible family, exposing outcomes that selected-branch experience leaves unobserved. Building on this, we propose Execution-Aware Learning (EAL), which combines Monte Carlo tree search with Q-learning to distill these outcomes into family-conditioned values. At deployment, the coordinator selects the policy family according to observable context, and the frozen coding agent generates the next local code block. Comprehensive single-episode evaluations on 100 tasks show that RoboAware reaches a 77.0% overall success rate, with SOTA averages of 90.0% on RoboSuite, 73.8% on diverse LIBERO-Pro task clusters, and 90.0% on challenging RoboTwin bimanual tasks, outperforming existing code-as-policy and VLA-harness baselines.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑