arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RoboHarn-Evo:演化分层物理知识以实现自改进的机器人操作

RoboHarn-Evo: Evolving Hierarchical Physical Knowledge for Self-Improving Robotic Manipulation

Shifeng Bao, Fanding Huang, Yihan Lin, Youhe Feng, Guanlin Li, Chen Zhao, Yang Li, Jiawei He, Cheng Chi, Jing Zhang

arXiv 2609.37583首次发表:更新:

发表机构

Renmin University of China; Tsinghua University; Key Laboratory of Data Engineering and Knowledge Engineering, Beijing, China; XYZ Embodied AI; Engineering Research Center of Database and Business Intelligence, Beijing, China; Beijing Academy of Artificial Intelligence (BAAI)(中国人民大学; 清华大学; 数据工程与知识工程重点实验室; XYZ具身智能公司; 数据库与商务智能工程研究中心; 北京智源人工智能研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出RoboHarn-Evo双循环框架,通过演化分层物理知识(任务知识与动作知识)利用物理反馈自改进机器人操作,在RMBench上平均成功率提升达24.2个百分点,并实现零样本迁移。

AI 中文摘要

视觉语言模型能够协调长时程机器人操作,然而成功的任务推理仍取决于局部物理交互是否产生预期效果。我们研究了在不更新基础模型的情况下,重复交互如何提升这种能力。我们引入了RoboHarn-Evo,一种双循环框架,从物理经验中演化分层物理知识(HPK)。HPK耦合了两层可复用知识:任务知识捕获应执行哪个子任务以及何时完成,而动作知识捕获对象相对几何策略及其物理效果。在执行过程中,智能体在相应的决策层级检索知识,并在任务目标下将其落地到当前场景。跨回合中,物理反馈用于修订历史知识、更新其适用性,并组织可复用条目以供后续检索。在RMBench上的实验表明,HPK在不同智能体模型上将平均成功率提升了高达24.2个百分点。通过80次交互 rollout,GPT-5.5的留出成功率从48.3%提升至75.0%,GPT-6从70.0%提升至88.3%。RoboHarn-Evo还解决了超过83%的历史知识错误,同时保留了95.8%的有效知识,并从RMBench零样本迁移至RoboDojo,分别获得35.0和25.0个百分点的提升。这些结果表明,物理交互可以累积为可复用知识,以改进后续操作。

英文摘要

Vision-language models can coordinate long-horizon robot manipulation, yet successful task reasoning still depends on whether local physical interactions produce the intended effects. We study how repeated interaction can improve this capability without updating the base model. We introduce RoboHarn-Evo, a dual-loop harness that evolves Hierarchical Physical Knowledge (HPK) from physical experience. HPK couples two levels of reusable knowledge: Task Knowledge captures which subtask should be executed and when it is complete, while Action Knowledge captures object-relative geometric strategies and their physical effects. During execution, the agent retrieves knowledge at the corresponding decision level and grounds it in the current scene under the task goal. Across episodes, physical feedback is used to revise historical knowledge, update its applicability, and organize reusable entries for subsequent retrieval. Experiments on RMBench show that HPK improves average success by up to 24.2 percentage points across different agent models. With 80 interaction rollouts, held-out success rises from 48.3% to 75.0% for GPT-5.5 and from 70.0% to 88.3% for GPT-6. RoboHarn-Evo also resolves over 83% of historical knowledge errors while retaining 95.8% of valid knowledge, and transfers zero-shot from RMBench to RoboDojo with gains of 35.0 and 25.0 percentage points. These results demonstrate that physical interaction can be accumulated into reusable knowledge for improving subsequent manipulation.

Comments39 pages, 9 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑