arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.01234cs.AI

在大语言模型自进化中通过自适应记忆-参数协调学习该记住什么和该内化什么

Learning What to Remember and What to Internalize in LLM Self-Evolution via Adaptive Memory-Parameter Coordination

  • University of Science and Technology of China(中国科学技术大学)
  • City University of Hong Kong(香港城市大学)
  • Hefei Normal University(合肥师范学院)
  • Xi’an Jiaotong University(西安交通大学)

机构由 AI 辅助整理,请以论文原文为准。

Tianyun Ji, Zhenya Huang, Jiayu Liu, Zirui Liu, Yu Su, Hongbin Pei

AI总结:

该研究提出 COVE 框架,通过任务感知路由等方式协调大语言模型自进化的记忆与参数通道,解决单通道进化的灵活性与性能权衡问题,在多任务上实现更稳健高效的改进。

AI中文摘要:

大语言模型智能体越来越多地在动态环境中运行,其中工具接口、API和用户需求在部署后会发生变化。现有的自进化方法主要遵循两种范式:基于工具 harness 的方法,将反馈外部化为可编辑的记忆或技能以实现快速适应;基于参数的方法,将经验内化到模型参数中以实现更深层次的能力提升。然而,单独使用任一机制都会在灵活性和性能之间产生权衡。本文研究智能体如何协调这两个通道以实现稳健的自进化,提出 COVE 这一统一的智能体自进化框架,通过任务感知路由、阶段感知调度和知识优化结合基于工具 harness 和基于参数的学习。通过该设计,COVE 将自进化视为任务与知识类型匹配适当学习机制的协调过程,而非经验的无差别积累。在多个任务类别上的实验表明,COVE 优于单通道进化策略,在变化环境下展现出更稳健、高效的改进。

英文摘要:

Large language model agents increasingly operate in dynamic environments where tool interfaces, APIs, and user requirements change after deployment. Existing self-evolution methods mainly follow two paradigms: harness-based approaches, which externalize feedback into editable memories or skills for rapid adaptation, and parameter-based approaches, which internalize experience into model parameters for deeper capability improvement. However, using either mechanism alone creates a trade-off between flexibility and performance. This paper asks how an agent can coordinate both channels to achieve robust self-evolution. We present COVE, a unified agent self-evolution framework that combines harness-based and parameter-based learning through task-aware routing, stage-aware scheduling, and knowledge optimization. Through this design, COVE treats self-evolution not as indiscriminate accumulation of experience, but as a coordinated process that matches tasks and knowledge types to appropriate learning mechanisms. Experiments across multiple task categories show that COVE outperforms single-channel evolution strategies, demonstrating more robust and efficient improvement under changing environments.

↑