arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.19013cs.LGcs.AI

利用工具链持续学习:超越模型参数的持续适应

Harness Continual Learning: Continual Adaptation Beyond Model Parameters

Borui Kang, Jinrui Gu, Junhan Lv, Wenbin Li, Lei Wang, Yang Gao

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出工具链持续学习(HCL)范式,围绕冻结基础模型演化工具链,通过受保护演化机制缓解工具链级遗忘,在多任务实验中相对基准增益超10%,可调整稳定性-可塑性权衡。

中文摘要 AI 辅助

持续学习在很大程度上是以模型为中心的,将模型参数视为随序列经验变化的状态。现代智能体还可以通过提示(prompt)、记忆、工具、技能和路由规则组成的工具链(harness)进行适应。由于这些内容共同塑造了后续的执行,即使模型被冻结,工具链的更新也可能破坏之前可靠的行为。这提出了一个新问题:智能体如何在模型之外持续改进其状态,同时保留早期获得的行为?我们提出了工具链持续学习(HCL)这一新的持续学习范式,其中工具链围绕冻结的基础模型演化,并将早期行为的损失定义为工具链级遗忘。我们用四个面向执行的组件实例化HCL:任务接口、经验记忆、能力映射和自适应路由。我们进一步引入受保护的工具链演化,将更新生成与状态承诺分离:持续优化器根据执行后反馈提出候选工具链,持续评估器仅在检查当前改进、历史保留情况和有效性后才承诺该候选工具链。在文本推理、多模态感知和开放世界交互上的实验表明,该方法具备能力积累和故障恢复能力,在多个设置中比相应基准的相对增益超过10%。组件消融实验评估了每个工具链组件的贡献,而受控保留扫描揭示了可测量的工具链级遗忘,并表明可以显式调整稳定性-可塑性权衡。

英文摘要

Continual learning has largely been model-centric, treating model parameters as the state that changes with sequential experience. Modern agents can also adapt through a harness of prompts, memories, tools, skills, and routing rules. Because these contents jointly shape later execution, a harness update can disrupt previously reliable behavior even when the model is frozen. This raises a new question: how can an agent continually improve its state outside the model while retaining behavior acquired earlier? We formulate Harness Continual Learning (HCL), a new continual learning paradigm in which the harness evolves around a frozen foundation model, and define the resulting loss of earlier behavior as harness-level forgetting. We instantiate HCL with four execution-facing components: the Task Interface, Experience Memory, Capability Map, and Adaptive Router. We further introduce guarded harness evolution to separate update generation from state commitment. A Continual Optimizer proposes candidate harnesses from post-execution feedback, and a Continual Evaluator commits the resulting candidate harness only after checking current improvement, historical retention, and validity. Experiments on textual reasoning, multimodal perception, and open-world interaction demonstrate capability accumulation and failure recovery, with relative gains exceeding 10% over corresponding baselines in multiple settings. Component ablations assess the contribution of each harness component, while controlled retention sweeps reveal measurable harness-level forgetting and show that the stability--plasticity trade-off can be explicitly adjusted.

发表机构

  • University of Wollongong(卧龙岗大学)
  • Nanjing University(南京大学)

机构由 AI 辅助整理,请以论文原文为准。

↑