AI 中文总结
RoboAware是一种基于反事实结果学习的具身技能协调方法,通过$P^5$架构、SCB和EAL技术,在100个任务上总体成功率达77.0%,优于现有基线方法。
AI 中文摘要
具身编码智能体可将模块化机器人技能与冻结的端到端策略相结合,但有效的组合需要预判当前物理状态下哪种策略族会成功。我们提出RoboAware,它在编码智能体的技能编排基础上,仅从反事实结果中学习状态条件下的责任协调器。受REPL成功的启发,我们提出$P^5$架构并基于其构建分层马尔可夫决策过程(MDP)。$P^5$将技能统一组织为五个语义阶段,定义了可比较责任的位置。为解决现有工作中缺乏反事实分支结果的问题,我们引入状态锁定反事实分支(SCB),该方法恢复相同的训练状态,以从每个可接受的策略族生成并执行代码块,从而揭示选定分支经验未观察到的结果。在此基础上,我们提出执行感知学习(EAL),它将蒙特卡洛树搜索与Q学习相结合,以将这些结果提炼为策略族条件值。部署时,协调器根据可观察上下文选择策略族,冻结的编码智能体生成下一个本地代码块。对100个任务的全面单回合评估显示,RoboAware的总体成功率达77.0%,在RoboSuite上达到90.0%的SOTA平均值、在多样化LIBERO-Pro任务集群上达到73.8%、在具有挑战性的RoboTwin双臂任务上达到90.0%,优于现有的代码即策略和VLA-harness基线方法。
英文摘要
Embodied coding agents can combine modular robot skills with frozen end-to-end policies, yet effective composition requires anticipating which policy family will succeed in the current physical state. We present RoboAware, which builds on coding agents' skill orchestration by learning only a state-conditioned responsibility coordinator from counterfactual outcomes. Inspired by the success of REPL, we propose the $P^5$ schema and formulate a hierarchical MDP based on it. $P^5$ organizes skills uniformly into five semantic stages, defining where responsibility can be compared. To address the lack of counterfactual branch outcomes in existing work, we introduce State-Locked Counterfactual Branching (SCB), which restores the same training state to generate and execute a code block from each admissible family, exposing outcomes that selected-branch experience leaves unobserved. Building on this, we propose Execution-Aware Learning (EAL), which combines Monte Carlo tree search with Q-learning to distill these outcomes into family-conditioned values. At deployment, the coordinator selects the policy family according to observable context, and the frozen coding agent generates the next local code block. Comprehensive single-episode evaluations on 100 tasks show that RoboAware reaches a 77.0% overall success rate, with SOTA averages of 90.0% on RoboSuite, 73.8% on diverse LIBERO-Pro task clusters, and 90.0% on challenging RoboTwin bimanual tasks, outperforming existing code-as-policy and VLA-harness baselines.