发表机构
Sakana AI(Sakana AI)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究将闭环机器人实现视为可复用执行经验,通过编码智能体迭代优化生成策略代码,实验表明优化参考代码在九任务平均成功率上较初始参考提升15.6个百分点,验证了其迁移价值。
AI 中文摘要
闭环机器人策略需要观测处理、状态管理和依赖情境的分支,这使得手动设计和调优成本高昂。尽管编码智能体日益支持控制代码的生成和优化,但尚不清楚在源任务上改进的实现是否也能支持新任务的策略获取。我们通过将完整的闭环实现视为可复用的执行经验来研究这一问题。对于每个源任务,编码智能体从少量成功演示中生成策略代码,并利用仿真反馈进行迭代改进。经过验证选择的实现被保留在软件档案中。对于新任务,智能体利用档案中的实现、目标演示和执行反馈来生成并改进策略。最终得到的策略被冻结,并在无需进一步模型调用的情况下执行。在RoboCasa的四个源任务中,迭代优化将平均成功率从28.3%提升至64.2%。在九个目标任务和三次独立运行中,无参考时的平均成功率为45.2%,使用初始源代码时为41.5%,使用优化源代码时为57.0%。在九任务平均值上,优化参考在所有三次运行中均优于初始参考,平均提升15.6个百分点。这些结果表明,在此设置中,经过执行改进的软件作为获取新策略的资源具有价值,尽管在跨运行平均时,初始参考在两个目标任务上仍然更优。
英文摘要
Closed-loop robot policies require observation processing, state management, and situation-dependent branching, making them costly to design and tune manually. Although coding agents increasingly support control-code generation and optimization, it remains unclear whether implementations improved on source tasks also support policy acquisition for new tasks. We study this question by treating complete closed-loop implementations as reusable execution experience. For each source task, a coding agent generates policy code from a few successful demonstrations and iteratively improves it using simulation feedback. The validation-selected implementations are retained in a software archive. For new tasks, the agent generates and improves policies using archived implementations, target demonstrations, and execution feedback. The resulting policy is then frozen and executes without further model calls. Across four source tasks in RoboCasa, iterative optimization increases mean success from 28.3% to 64.2%. Across nine target tasks and three independent runs, mean success is 45.2% without references, 41.5% with initial source code, and 57.0% with optimized source code. Optimized references outperform initial references in all three runs on the nine-task average, with a mean gain of 15.6 percentage points. These results demonstrate the value of execution-improved software as a resource for acquiring new policies in this setting, although initial references remain better on two target tasks when averaged across runs.