arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.23989cs.AI

ACLArena:智能体在多阶段后训练中的持续学习

ACLArena: Agent Continue Learning in Multi-stage Post-training

Haixin Wang, Xiaoxuan Wang, Junkai Zhang, Han Zhang, Renliang Sun, Alexander K Taylor, Yidan Shi, Haoran Deng, Chenguang Wang, Jason Cong, Yizhou Sun, Wei Wang

首次发表
浏览论文内容

中文总结 AI 辅助

针对智能体多阶段训练中能力整合缺乏成熟方案的问题,提出ACLArena框架,通过分析遗忘与泛化机制,结合离线回放与多LoRA专家路由网络,显著提升多领域持续学习能力。

中文摘要 AI 辅助

构建用于工业部署的通用智能体需要整合多种能力,每种能力通常在训练的不同阶段获得。然而,目前尚无成熟的智能体持续学习(ACL)方案,对现有整合范式之间的权衡也知之甚少。为填补这一空白,我们提出了ACLArena,一个用于全面研究、分析和评估ACL的框架。我们首先构建了一个顺序训练流程,并从模型层面和令牌层面两个互补视角进行了深入分析,以解释遗忘和泛化的机制。在这些分析的指导下,我们系统地比较了多教师在线蒸馏、自蒸馏微调和模型合并,以评估它们在保留新获得能力的同时恢复先前学习能力的效果。通过大量实验,我们深入理解了能力如何在各阶段之间迁移。最后,我们提出了一种新的ACL方案,该方案将高质量轨迹的离线回放与多个通过强化学习专门化的LoRA专家的路由网络相结合,显著提升了智能体在多个领域学习的能力。在四个推理和智能体任务上进行的综合实验,在域内和域外设置下均进行了评估,证明了我们分析的价值和方法的有效性。

英文摘要

Building general-purpose agents for industrial deployment requires integrating multiple capabilities, each typically acquired at a distinct stage of training. Yet there is currently no well-established recipe for Agent Continual Learning (ACL), with little understanding of the trade-offs among existing integration paradigms. To address this gap, we introduce ACLArena, a framework for comprehensively studying, analyzing, and evaluating ACL. We first build a sequential training pipeline and conduct an in-depth analysis that explains the mechanisms of forgetting and generalization from two complementary perspectives, the model level and the token level. Guided by these analyses, we systematically compare multi-teacher on-policy distillation, self-distilled fine-tuning, and model merging to assess their ability to recover previously learned capabilities while preserving newly acquired ones. Through extensive experiments, we develop a detailed understanding of how capabilities transfer across stages. Finally, we propose a new ACL recipe that combines offline replay over high-quality trajectories with a routed network of multiple LoRA experts each specialized via RL, substantially improving the agent's ability to learn across multiple domains. Comprehensive experiments on four reasoning and agentic tasks, evaluated under both in-domain and out-of-domain settings, demonstrate the value of our analysis and the effectiveness of our approach.

补充信息

↑