发表机构
The Hong Kong Polytechnic University; Huawei; Renmin University of China(香港理工大学; 华为; 中国人民大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
经验漏斗通过状态与策略交替循环,将交互经验先提炼为显式状态快速适应,再蒸馏为策略参数,实现智能体自进化并提升能力。
AI 中文摘要
由大型语言模型(LLMs)驱动的自主智能体通过交互不断积累经验,这为通过自我进化来改进未来行为创造了机会。一个根本性的挑战是如何将丰富的、特定于任务的交互经验转化为可复用的模型能力,同时不牺牲快速适应新观察证据的能力。显式的文本状态(如技能和智能体框架)提供了快速、人类可读且可编辑的适应方式,但会导致对外部上下文的持续依赖;参数化策略提供了紧凑且可复用的能力,但更新速度要慢得多。我们提出了经验漏斗(Experience Funnel),一种自进化框架,它以交替循环的方式将快速状态适应与缓慢策略整合相结合。交互轨迹首先被提炼为显式文本状态,新获得的经验可以在此快速纳入并验证。该框架随后有选择地识别在状态修订中仍然有用的状态启用的行为,并通过过渡感知蒸馏将其整合到策略中。更新后的状态-策略对随后生成新的轨迹,为下一轮状态适应和策略整合提供新的证据。在多个智能体基准上的实验表明,经验漏斗(Experience Funnel)在智能体能力上持续优于仅状态进化和策略内化方法,同时逐步将有用的显式经验转化为自主策略能力。
英文摘要
Autonomous agents powered by large language models (LLMs) continuously accumulate experience through interaction, creating an opportunity to improve future behavior through self-evolution. A fundamental challenge is how to transform abundant, task-specific interaction experience into reusable model competence without sacrificing the ability to adapt rapidly to newly observed evidence. Explicit textual states, such as skills and agent harnesses, provide fast, human-readable and editable adaptation, but incur persistent dependence on external context; parametric policies provide compact and reusable competence, but are substantially slower to update. We present \textit{Experience Funnel}, a self-evolving framework that couples fast state adaptation with slow policy consolidation in an alternating loop. Interaction trajectories are first distilled into an explicit textual state, where newly acquired experience can be rapidly incorporated and validated. The framework then selectively identifies state-enabled behavior that remains useful across state revisions and consolidates it into the policy through transition-aware distillation. The updated state--policy pair subsequently generates new rollouts, providing fresh evidence for the next round of state adaptation and policy consolidation. Experiments across diverse agent benchmarks show that \textit{Experience Funnel} consistently improves agent capability over state-only evolution and policy-internalization approaches, while progressively converting useful explicit experience into autonomous policy competence.