EMERGE-Policy:超越单一策略的机器人心智涌现
EMERGE-Policy: A Robot Mind Emerges Beyond a Single Policy
浏览论文内容
中文总结 AI 辅助
EMERGE-Policy是一种图结构智能体框架,通过多角色子智能体协作划分功能子任务,无需额外微调即可在公开基准和真实机器人实验中实现出色性能,扩展了机器人策略的应用范围。
中文摘要 AI 辅助
机器人的有效“心智”无需驻留在单一策略中,它可在共享编排流程内,由专用组件执行感知、推理、预测、行动、验证与记忆的过程中涌现。EMERGE-Policy将该视角转化为图结构智能体框架,协调能力调用与信息交换:主智能体在活跃上下文窗口中保留任务级状态;角色专用子智能体在独立上下文中处理感知、执行监控、验证与记忆整合,并返回结构化的任务相关证据;角色专用上下文通过仅向主智能体暴露决策相关证据控制信息负载;功能性Skill接口将异构后端组合为操作Skill、想象Skill与评估Skill;基于准则的验证、文本故障诊断及分支栈恢复提供局部修正,令牌感知外部记忆保留任务相关状态。其闭环交互共同实现了EMERGE-Policy所代表的系统级策略。在无需额外微调的情况下,该框架在多个具有广泛影响力的公开基准上取得了出色性能,并开展了一系列真实机器人实验。这些系统级结果表明,通过在多个智能体间划分不同功能子任务并使其并行协作,以及将模型视为框架内可调用Skill的技术范式,EMERGE-Policy可将稳健的机器人策略扩展至孤立运行之外。
英文摘要
A robot's effective ``mind'' need not reside in a single policy. It can emerge when specialized components perceive, reason, predict, act, verify, and remember within a shared orchestration process. EMERGE-Policy turns this perspective into a graph-structured agentic framework that coordinates both capability invocation and information exchange. A Main Agent retains task-level state within an active context window, while role-specific Sub Agents process perception, execution monitoring, verification, and memory consolidation in isolated contexts and return structured, task-relevant evidence. Role-specific contexts control information load by exposing only decision-relevant evidence to the Main Agent, while the functional Skill interface composes heterogeneous backends as Operational, Imagination, and Evaluation Skills. Criterion-grounded verification, textual failure diagnosis, and Branch Stack recovery provide localized correction, with token-aware external memory preserving task-relevant state. Together, their closed-loop interaction realizes the system-level policy captured by the name EMERGE-Policy. Without additional fine-tuning, we achieved outstanding performance on several public benchmark that have had a wide-reaching impact, and conducted a series of real robot experiments. These system-level results suggest that through the division of different functional sub-tasks among multiple agents and their concurrent collaboration, as well as the technical paradigm where the model is regarded as a skill and called within the framework, EMERGE-Policy can extend the robust robot policies beyond isolated runs.
发表机构
- Tsinghua University(清华大学)
- Nanjing University of Science and Technology(南京理工大学)
- Xi’an Jiaotong University(西安交通大学)
- Xidian University(西安电子科技大学)
- Harbin Institute of Technology(哈尔滨工业大学)
- Peking University(北京大学)
- Nanyang Technological University(南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。