arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于技能-工具链进化的自演化具身智能体

Self-Evolving Embodied Agents via Skill-Harness Evolution

Peidong Wang, Zhiming Ma, Ying Chang, Xufang Luo, Yiqun Zhang, Zihan Wang, Xiaocui Yang, Shi Feng, Yuqing Yang, Dongsheng Li

arXiv 2608.11350首次发表:更新:

发表机构

Microsoft Research; Northeastern University(微软研究院; 东北大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出免训练的SHAPER框架,通过演化技能与工具链实现冻结模型驱动的自演化具身智能体,在VLABench等数据集上验证其适配优势。

AI 中文摘要

具身智能体越来越多地围绕基础模型构建,其性能不仅取决于模型权重,还围绕模型的技能、上下文、动作接口和执行工具链。尽管监督微调与强化学习可使智能体适应新环境,但它们需要额外数据、奖励及训练轮次;与此同时,许多免训练的代码中心方法依赖可编程机器人API,而这些API在固定接口场景中可能无法使用。我们提出SHAPER,一种免训练的具身适配自演化框架,该框架保持模型参数冻结,通过目标环境的 rollout(试跑)演化可复用技能与上下文-代码工具链,以此改进非参数化智能体系统。在SHAPER中,同一冻结模型可同时充当规划器与优化器,无需参数更新即可优化其外部技能及上下文-代码工具链。我们在涵盖不同低级动作接口具身智能体的VLABench与ESI-Bench上评估SHAPER,并将其与纯执行、监督微调及无验证器选择、投票等测试时缩放基线对比。结果表明,当模型训练成本高、不可用或不合意时,技能与工具链优化是实现自演化具身智能体的可行路径。

英文摘要

Embodied agents are increasingly built as systems around foundation models, where performance depends not only on model weights but also on the skills, context, action interfaces, and execution harness surrounding the model. While supervised fine-tuning and reinforcement learning can adapt agents to new environments, they require additional data, rewards, and training runs; meanwhile, many train-free code-centric approaches rely on programmable robot APIs that may be unavailable in fixed-interface settings. We propose SHAPER, a self-evolving framework for train-free embodied adaptation that keeps model parameters frozen and improves the non-parametric agent system by evolving reusable skills and a context-code harness through target-environment rollouts. In SHAPER, the same frozen model can serve as both planner and optimizer, refining its external skills and context-code harness without parameter updates. We evaluate SHAPER on VLABench and ESI-Bench, covering embodied agents with different low-level action interfaces, and compare against pure execution, supervised fine-tuning, and test-time-scaling baselines such as verifier-free selection and voting. Our results suggest that skill-and-harness optimization is a practical route to self-evolving embodied agents when model training is expensive, unavailable, or undesirable.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑