arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.29896cs.RO

EMERGE-Policy:超越单一策略的机器人心智涌现

EMERGE-Policy: A Robot Mind Emerges Beyond a Single Policy

Zhirui Fang, Qingchi Yu, Ziyang Chen, Longfei Li, Haoran Ma, Keru Zhou, Xinrun Xu, Samith Va, Yuxuan Hu, Peixuan Song, Qiang Du, Bin Qian, Yongkang Deng, Xin Li… 展开作者

Zhirui Fang, Qingchi Yu, Ziyang Chen, Longfei Li, Haoran Ma, Keru Zhou, Xinrun Xu, Samith Va, Yuxuan Hu, Peixuan Song, Qiang Du, Bin Qian, Yongkang Deng, Xin Li, Yezhen Wang, Zhe Li, Hao Luo, Shuyan Li, Ziwei Wang, Weijian Deng, Xiu Li

首次发表
浏览论文内容

中文总结 AI 辅助

EMERGE-Policy是一种图结构智能体框架,通过多角色子智能体协作划分功能子任务,无需额外微调即可在公开基准和真实机器人实验中实现出色性能,扩展了机器人策略的应用范围。

中文摘要 AI 辅助

机器人的有效“心智”无需驻留在单一策略中,它可在共享编排流程内,由专用组件执行感知、推理、预测、行动、验证与记忆的过程中涌现。EMERGE-Policy将该视角转化为图结构智能体框架,协调能力调用与信息交换:主智能体在活跃上下文窗口中保留任务级状态;角色专用子智能体在独立上下文中处理感知、执行监控、验证与记忆整合,并返回结构化的任务相关证据;角色专用上下文通过仅向主智能体暴露决策相关证据控制信息负载;功能性Skill接口将异构后端组合为操作Skill、想象Skill与评估Skill;基于准则的验证、文本故障诊断及分支栈恢复提供局部修正,令牌感知外部记忆保留任务相关状态。其闭环交互共同实现了EMERGE-Policy所代表的系统级策略。在无需额外微调的情况下,该框架在多个具有广泛影响力的公开基准上取得了出色性能,并开展了一系列真实机器人实验。这些系统级结果表明,通过在多个智能体间划分不同功能子任务并使其并行协作,以及将模型视为框架内可调用Skill的技术范式,EMERGE-Policy可将稳健的机器人策略扩展至孤立运行之外。

英文摘要

A robot's effective ``mind'' need not reside in a single policy. It can emerge when specialized components perceive, reason, predict, act, verify, and remember within a shared orchestration process. EMERGE-Policy turns this perspective into a graph-structured agentic framework that coordinates both capability invocation and information exchange. A Main Agent retains task-level state within an active context window, while role-specific Sub Agents process perception, execution monitoring, verification, and memory consolidation in isolated contexts and return structured, task-relevant evidence. Role-specific contexts control information load by exposing only decision-relevant evidence to the Main Agent, while the functional Skill interface composes heterogeneous backends as Operational, Imagination, and Evaluation Skills. Criterion-grounded verification, textual failure diagnosis, and Branch Stack recovery provide localized correction, with token-aware external memory preserving task-relevant state. Together, their closed-loop interaction realizes the system-level policy captured by the name EMERGE-Policy. Without additional fine-tuning, we achieved outstanding performance on several public benchmark that have had a wide-reaching impact, and conducted a series of real robot experiments. These system-level results suggest that through the division of different functional sub-tasks among multiple agents and their concurrent collaboration, as well as the technical paradigm where the model is regarded as a skill and called within the framework, EMERGE-Policy can extend the robust robot policies beyond isolated runs.

发表机构

  • Tsinghua University(清华大学)
  • Nanjing University of Science and Technology(南京理工大学)
  • Xi’an Jiaotong University(西安交通大学)
  • Xidian University(西安电子科技大学)
  • Harbin Institute of Technology(哈尔滨工业大学)
  • Peking University(北京大学)
  • Nanyang Technological University(南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

↑