arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.14633cs.RO

REVOLVE:一种最小人工干预下机器人操作演化的自动化闭环框架

REVOLVE: An Automated Closed-Loop Framework for Evolving Robot Manipulation with Minimal Human Intervention

Hanyu Liu, Qian Li, Yizhu Ding, Jiayi Wen, Keqiang Ren, Yunsheng Ma, Tao Jian, Zhihua Wang, Zhuofan Yu, Xinran Li, Zhigong Song

首次发表
浏览论文内容

中文总结 AI 辅助

REVOLVE提出自动化闭环框架,通过自动重置与收集及双循环演化,在最小人工干预下持续提升机器人操作策略和智能体,实验显示成功率提升18.5%,人工工作量大幅减少。

中文摘要 AI 辅助

近年来,数据驱动的机器人操作策略在任务执行和泛化方面取得了显著进步。然而,实际部署仍然严重依赖人工进行失败评估、纠正和环境重置,而模型往往无法从失败和纠正经验中持续学习。我们提出了REVOLVE(通过编排循环、验证和经验进行机器人演化),一种自动化闭环框架,用于在最小人工干预下演化机器人操作。基于统一的软件平台,REVOLVE将数据收集、策略训练与部署、失败恢复和持续学习集成到单一的闭环工作流中。其自动重置与收集(ARC)架构自动重置环境并干预以纠正策略失败。双循环演化(DLE)通过将真实世界交互和失败-纠正数据反馈到策略学习中,并使用外部不匹配记忆来细化智能体判断,从而持续改进操作策略和智能体。在四个真实世界操作任务上的实验表明,经过五次迭代后,REVOLVE将平均策略成功率提高了18.5%,智能体判断准确率提高了8.5%,同时将数据收集和部署测试中的人工工作量分别减少了94.4%和95.1%。这些结果表明,REVOLVE将真实世界部署转变为一种闭环学习过程,持续积累和使用执行经验,使得策略和监督模型能够在显著减少人工干预的情况下持续演化。

英文摘要

Recent advances in data-driven robot manipulation policies have substantially improved task execution and generalization. However, real-world deployment still relies heavily on humans for failure assessment, correction, and environment reset, while models often fail to continually learn from failures and corrective experience. We present REVOLVE (Robot Evolving via Orchestrated Loops, Verification, and Experience), an automated closed-loop framework for evolving robot manipulation with minimal human intervention. Built on a unified software platform, REVOLVE integrates data collection, policy training and deployment, failure recovery, and continual learning into a single closed-loop workflow. Its Automated Reset and Correction (ARC) architecture automatically resets the environment and intervenes to correct policy failures. Dual-Loop Evolution (DLE) continually improves the manipulation policy and agent by feeding real-world interaction and failure--correction data back into policy learning and using an external mismatch memory to refine agent judgments. Experiments across four real-world manipulation tasks show that, after five iterations, REVOLVE improves average policy success rate by 18.5% and agent judgment accuracy by 8.5%, while reducing human effort in data collection and deployment testing by 94.4% and 95.1%, respectively. These results demonstrate that REVOLVE transforms real-world deployment into a closed-loop learning process that continually accumulates and uses execution experience, enabling continual evolution of both the policy and supervisory model with substantially less human intervention.

发表机构

  • Jiangnan University(江南大学)

机构由 AI 辅助整理,请以论文原文为准。

↑