RoboFoundry:面向自学习具身智能体的系统即策略演化
RoboFoundry: System-as-Policy Evolution for Self-Learning Embodied Agents
浏览论文内容
中文总结 AI 辅助
RoboFoundry提出系统即策略的演化框架,将执行经验转化为系统更新,在多个基准上显著提升性能,实现跨机器人迁移与自主演化。
中文摘要 AI 辅助
基础模型不应作为具身智能体孤立地行动。然而,现有方法往往优化智能体栈的各个独立组件,如记忆、上下文、技能或动作接口,而非将支撑系统本身视为统一的策略。此外,仅靠交互并不能带来自我改进,除非执行经验被转化为持久且经过验证的系统变更。因此,我们提出RoboFoundry,这是首个将这一过程形式化为自演化系统即策略的具身智能体框架。RoboFoundry诊断决策和记忆管理中的能力缺口,将执行轨迹转化为经过验证的特定任务系统更新,并促进对通用系统的持续改进。演化在两个互补的层面上进行:一个管理活动内部上下文和持久文件系统记忆的上下文系统,以及一个组织原子技能、可复用组合和失败条件恢复的分层技能系统。一个共享的语义接口将不依赖具身形态的决策与依赖具身形态的执行分离,使得演化后的系统能力能够在异构机器人之间迁移。在EmbodiedBench上,RoboFoundry达到了最先进的性能,显著提升了GPT-5.5达27.8%。它还使Qwen3.7-Plus接近GPT-5.5的水平(70.3%对72.7%),显示出系统即策略演化在不同基础模型上的一致增益。对于长时程记忆,RoboFoundry在RoboMemArena上至少超过所有基线39.0%,甚至优于有外部基础模型辅助的方法。在LIBERO-PRO上,它在所有扰动类型上进一步超过Cap-Agent0达243.8%-679.7%。在真实世界部署中,RoboFoundry展示了跨机器人和任务的零样本迁移和在线演化,突显了其实现完全自主具身智能体的潜力。
英文摘要
A foundation model should not act in isolation as an embodied agent. Yet, existing methods often optimize individual components of the agent stack, such as memory, context, skills, or action interfaces, rather than treating the supporting system itself as a unified policy. Moreover, interaction alone does not yield self-improvement unless execution experience is converted into persistent, validated system changes. We therefore propose RoboFoundry, the first embodied agentic framework that formulates this process as Self-Evolving System-as-Policy. RoboFoundry diagnoses capability gaps in decision-making and memory management, converts execution traces into validated task-specific system updates, and promotes recurring improvements to the general system. Evolution operates over two complementary surfaces: a context system that manages active internal context and persistent file-system memory, and a hierarchical skill system that organizes atomic skills, reusable compositions, and failure-conditioned recovery. A shared semantic interface separates embodiment-invariant decisions from embodiment-specific execution, allowing evolved system capabilities to transfer across heterogeneous robots. On EmbodiedBench, RoboFoundry achieves state-of-the-art performance, notably improving GPT-5.5 by 27.8%. It also brings Qwen3.7-Plus to near parity with GPT-5.5 (70.3% vs. 72.7%), showing consistent gains from system-as-policy evolution across foundation models. For long-horizon memory, RoboFoundry outperforms all baselines on RoboMemArena by at least 39.0%, even against methods assisted by external foundation models. On LIBERO-PRO, it further outperforms Cap-Agent0 by 243.8%-679.7% across all perturbation types. In real-world deployments, RoboFoundry demonstrates zero-shot transfer and online evolution across robots and tasks, highlighting its potential for fully autonomous embodied agents.