发表机构
Beijing Academy of Artificial Intelligence (BAAI); Institute of Software, Chinese Academy of Sciences; University of Chinese Academy of Sciences(北京人工智能研究院(BAAI); 中国科学院软件研究所; 中国科学院大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
PRACTICE通过训练技能学习者维护技能库,结合两阶段课程训练与在线编辑蒸馏,在EB-ALFRED等任务上实现优于基线的自进化具身智能体性能提升。
AI 中文摘要
近期研究表明,多模态大语言模型(MLLMs)可作为具身智能体,将语言指令与视觉观测转化为可执行计划。然而,构建能通过交互持续改进并快速适应环境的智能体仍具挑战性。总结过往交互轨迹的经验是有前景的解决方案,但现有基于经验的方法常依赖人工设计的提示工作流来提取和更新技能,这类固定流程难以从新的多样经验中学习更新后的技能。本文提出PRACTICE,其训练一个技能学习者,从过往交互轨迹中发现并维护持久的技能库,同时保持任务执行器冻结。给定历史累积技能和输入轨迹,技能学习者生成结构化的批量编辑,用于添加、优化、合并或移除技能,随后分层整合所有收集到的编辑,形成一致的更新后技能库。我们采用两阶段课程训练该学习者:首先,其从神示轨迹中学习基础技能生成与库维护;接着,通过对比相同任务下异构执行器的成功与失败轨迹,学习识别无效动作模式与恢复策略;最后,我们应用在线技能编辑蒸馏,使技能学习者在当前编辑分布上与更强的教师对齐,进一步优化策略。实验表明,紧凑的技能学习者在多轮连续库更新中,对多个冻结执行器实现了一致的性能提升;在EB-ALFRED和EB-Habitat任务上,PRACTICE进一步优于最强的基于经验的基线方法。项目资源可在this https URL获取。
英文摘要
Recent studies have shown that multimodal large language models (MLLMs) can serve as embodied agents, translating language instructions and visual observations into executable plans. However, building agents that can continually improve through interaction and rapidly adapt to their environments remains challenging. Summing up experience from past interaction trajectories provides a promising solution, but existing experience-based methods often rely on manually designed prompting workflows to extract and update skills. Such fixed procedures may struggle to learn updated skills from new and diverse experiences. We introduce PRACTICE, which trains a skill learner to discover and maintain a persistent skill library from past interaction trajectories while keeping the task executor frozen. Given the historical accumulated skills and incoming trajectories, the skill learner produces structured batch-edits that add, refine, merge, or remove skills, and then hierarchical consolidate all collected edits into a consistent updated skill library. We train the learner with a two-stage curriculum. First, it learns basic skill generation and library maintenance from oracle trajectories. Then, by contrasting successful and failed trajectories from heterogeneous executors on the same tasks, it learn to identify invalid action patterns and recovery strategies. Finally, we apply online skill-edit distillation to align the skill learner with a stronger teacher on its current edit distribution to further improves the policy. Experiments demonstrate that a compact skill learner delivers consistent performance improvements across successive library-update rounds for multiple frozen executors. On EB-ALFRED and EB-Habitat, PRACTICE further outperforms the strongest experience-based baselines. Project resources are publicly available at: https://baai-agents.github.io/PRACTICE