AI 中文总结
EmbodiRSI通过真实-仿真-真实循环中的递归自改进,利用策略反馈和两种互补机制(协同错误修正与自适应数据收集),在少量真实数据下将桌面操作成功率提升至83.1%,实现数据高效的具身策略适应。
AI 中文摘要
将机器人操作策略适应到新任务和环境仍然高度依赖数据,而进一步改进所需的数据取决于策略当前的能力和失败模式。我们提出了EmbodiRSI,一个在真实-仿真-真实(real-to-sim-to-real)设置中实现递归自改进(RSI)的智能体系统,其中从目标部署场景构建任务特定仿真,并将其用作在迁移回物理世界之前进行迭代策略改进的低成本环境。EmbodiRSI利用策略执行反馈来指导后续的经验获取和策略更新。两种互补机制闭环了该循环:协同错误修正(Collaborative Error Correction)从策略到达的状态生成智能体辅助的修正轨迹,而自适应数据收集(Adaptive Data Collection)将专家演示生成导向当前策略的弱点。任务特定仿真作为可复用工作空间,用于策略预热、可重复评估、失败诊断以及跨连续RSI轮次的目标数据生成。在三个桌面环境和14个子任务中,EmbodiRSI通过两次RSI更新将场景平衡的自主仿真成功率从50.4%提升至83.5%。每个子任务仅使用400条自适应仿真轨迹和十条真实世界细化轨迹,EmbodiRSI实现了83.1%的场景平衡自主真实世界成功率,而使用每个子任务200条真实世界演示进行适应的方法成功率为75.0%。这些结果表明,在部署特定仿真中进行反馈驱动的递归改进能够实现具身策略对物理环境的数据高效适应。
英文摘要
Adapting robot manipulation policies to new tasks and environments remains highly data-intensive, while the data needed for further improvement depends on the policy's current capabilities and failure modes. We introduce EmbodiRSI, an agentic system for recursive self-improvement (RSI) in a real-to-sim-to-real setting, where task-specific simulations are constructed from target deployment scenarios and used as low-cost environments for iterative policy improvement before transfer back to the physical world. EmbodiRSI uses policy execution feedback to guide subsequent experience acquisition and policy updates. Two complementary mechanisms close this loop: Collaborative Error Correction generates agent-assisted corrective trajectories from policy-reached states, while Adaptive Data Collection directs expert demonstration generation toward the current policy's weaknesses. The task-specific simulation serves as a reusable workspace for policy warm-up, repeatable evaluation, failure diagnosis, and targeted data generation across successive RSI rounds. Across three tabletop environments and 14 subtasks, EmbodiRSI increases scene-balanced autonomous simulation success from 50.4% to 83.5% over two RSI updates. With 400 adaptive simulated trajectories and only ten real-world refinement trajectories per subtask, EmbodiRSI achieves 83.1% scene-balanced autonomous real-world success, compared with 75.0% for adaptation using 200 real-world demonstrations per subtask. These results demonstrate that feedback-driven recursive improvement in deployment-specific simulations can enable data-efficient adaptation of embodied policies to physical environments.