RACaP:面向可进化机器人学习的智能体推理、行动与编码策略
RACaP: Agentic Reasoning, Acting, and Coding as Policies for Evolvable Robot Learning
浏览论文内容
中文总结 AI 辅助
RACaP将代码生成移至进化阶段,用ReAct循环调用冻结的策略API,在多个LIBERO基准上显著超越CaP基线,并通过蒸馏实现高效部署。
中文摘要 AI 辅助
通用型机器人智能体必须从经验中学习、迁移到新任务并高效行动。代码即策略(CaP)方法在运行时生成和修复程序,这会带来延迟,并将可复用机制与任务特定决策纠缠在一起。我们提出RACaP,一个将编码移至进化阶段并在部署时使用推理与行动(ReAct)循环调用冻结的、带类型的策略API的智能体框架。一种两阶段策略将能力课程学习与自主自我进化相结合,以改进API、ReAct框架和经验记忆。API编码可复用的物理机制,同时暴露用于运行时适应的参数。ReAct结合任务特定工作记忆、长期经验记忆和视觉反馈来选择动作、验证结果并从失败中恢复,而无需修改源代码。RACaP在LIBERO-90上达到54.4%的成功率,在零样本LIBERO-PRO上达到45.0%,在LIBERO-Long上达到46.0%,而CaP基线在长时程任务上最多仅为4.0%。在LIBERO-PRO上,其成功率是CaP基线的2.5倍,中位策略时间提速1.9倍。为实现高效的机载部署,拒绝采样微调将GPT-5.6的ReAct决策蒸馏到Qwen3-VL-8B-Instruct中,实现每次决策推理13.2倍的加速,并将重复物理调用从16次减少到4次。这些结果表明,将可复用代码与运行时决策分离支持持续进化、有效迁移和高效的长时程控制。
英文摘要
General-purpose robot agents must learn from experience, transfer to new tasks, and act efficiently. Code as Policies (CaP) methods generate and repair programs at runtime, incurring latency and entangling reusable mechanisms with task-specific decisions. We introduce RACaP, an agentic framework that moves coding to evolution and uses a Reasoning-and-Acting (ReAct) loop to call frozen, typed Policy APIs at deployment. A two-phase strategy combines capability curriculum learning with autonomous self-evolution to improve the APIs, the ReAct harness, and experience memory. The APIs encode reusable physical mechanisms while exposing arguments for runtime adaptation. ReAct combines task-specific working memory, long-term experience memory, and visual feedback to select actions, verify outcomes, and recover from failures without modifying source code. RACaP achieves 54.4% success on LIBERO-90, 45.0% on zero-shot LIBERO-PRO, and 46.0% on LIBERO-Long, compared with at most 4.0% for CaP baselines on long-horizon tasks. On LIBERO-PRO, it achieves 2.5 times the success rate of CaP baselines and a 1.9-fold speedup in median policy time. For efficient on-robot deployment, rejection-sampled fine-tuning distills GPT-5.6 ReAct decisions into Qwen3-VL-8B-Instruct, yielding a 13.2-fold per-decision inference speedup and reducing repeated physical calls from 16 to 4. These results show that separating reusable code from runtime decisions supports continued evolution, effective transfer, and efficient long-horizon control.
发表机构
- The Chinese University of Hong Kong(香港中文大学)
- The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
- Knowin AI(考恩人工智能)
机构由 AI 辅助整理,请以论文原文为准。