arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

从提示到功能库:演化功能库实现大语言模型的持续学习

From a Prompt to Repertoires: Evolving Functional REpertoires Enable LLM Continual Learning

Fengyuan Liu, Yue Wang, Hangxi Guo, Fengyuan Liu, Chenxu Wu, Yanguang Liu, Mengnan Du

arXiv 2610.11373首次发表:更新:

发表机构

The Chinese University of Hong Kong, Shenzhen; Shanghai AI Laboratory; University of Science and Technology of China; New Jersey Institute of Technology(香港中文大学(深圳); 上海人工智能实验室; 中国科学技术大学; 新泽西理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对大语言模型持续学习的灾难性遗忘问题,提出演化功能库(EFRE)方法,在三任务持续学习流上性能优于GRPO,且在智能体系统中适配不同主干模型均有提升。

AI 中文摘要

持续学习对于大语言模型而言仍是一项挑战,这类模型必须在获取新技能与知识的同时,不降低现有能力。现有方法通常通过精心设计模型参数的更新方式来应对这一挑战。相比之下,提示优化可避免代价高昂的参数更新,在知识密集型和推理类单个任务上,其性能甚至能达到或优于GRPO这类强化学习方法。这引出了一个自然的问题:作为一种高效的适配方法,提示优化能否直接应用于持续学习?我们的分析表明,在顺序任务适配下,提示优化会遭受灾难性遗忘,且优化后的提示会积累过拟合于局部任务分布的规则。为解决这些局限,我们提出了演化功能库(Evolving Functional REpertoires,EFRE),它用一个函数库替代单一提示,该函数库会随新任务的到来而演化:兼容的更新会优化现有函数,而冲突的更新则会触发新函数的出现。在包含三个任务的持续学习流上,EFRE的最终平均性能比GRPO高出7.50个百分点;此外,在适配Bio任务后,其在FinQA任务上的性能仅下降1.56个百分点,而基础提示优化方法的该指标下降了25.10个百分点。我们进一步在一个极简智能体系统中实例化了EFRE,观察到其在不同主干模型上均实现了一致的性能提升。总体而言,这些结果证明了EFRE在大语言模型持续学习中的优异性能,并凸显了其在高级智能体系统持续学习中的巨大潜力。

英文摘要

Continual learning remains challenging for large language models, which must enable models to acquire new skills and knowledge without degrading existing capabilities. Existing approaches typically address this challenge by carefully designing how model parameters are updated. In contrast, prompt optimization avoids costly parameter updates while achieving competitive or even superior performance to reinforcement learning methods such as GRPO on individual knowledge-intensive and reasoning tasks. This raises a natural question: \textit{Can prompt optimization, as an efficient adaptation approach, be directly applied to continual learning?} Our analysis shows that, under sequential task adaptation, it suffers from catastrophic forgetting, while optimized prompts accumulate rules that overfit to local task distributions. To address these limitations, we propose \emph{Evolving Functional REpertoires} (EFRE), which replaces a single prompt with a repertoire of functions that evolves as new tasks arrive: compatible updates refine existing functions, while conflicting updates trigger the emergence of new ones. On a three-task continual-learning stream, EFRE achieves a final average performance 7.50 percentage points higher than GRPO. Moreover, after adaptation to the Bio task, its performance on FinQA decreases by only 1.56 percentage points, compared with 25.10 percentage points for the base prompt optimization method. We further instantiate EFRE in a minimal agent system and observe consistent improvements across different backbone models. Overall, these results demonstrate EFRE's strong performance in continual learning for large language models and highlight its substantial potential for continual learning in advanced agent systems.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑