arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33398cs.AI

COEVO:递归自我改进的上下文与参数协同进化

COEVO: Co-Evolving Context and Parameters for Recursive Self-Improvement

Siwei Chen, Xinping Bao, Xinyu Cai, Yuan Cao, Wan Jiang, Shaohong Chen

首次发表
浏览论文内容

中文总结 AI 辅助

COEVO提出参数与上下文协同进化的框架,通过策略熵和注意力信号指导自适应,在递归自我改进中提升任务性能与鲁棒性。

中文摘要 AI 辅助

递归自我改进(RSI)旨在将大型语言模型从静态训练流程推向能够参与改进自身未来行为的系统。现有方法主要沿两个方向展开:通过在线学习更新模型参数,或通过搜索、反思和提示优化来改进外部上下文。尽管这两种机制都能支持持续改进,但它们通常被独立研究。这种分离忽略了一个重要的交互:上下文塑造了模型从中学习的经验,而不断进化的模型可能随时间以不同方式解释和利用相同的上下文。因此,我们将RSI表述为参数-上下文协同进化问题,其中模型参数和学习上下文在共享反馈回路中自适应。我们引入了COEVO框架,该框架从策略内经验中更新模型参数,同时根据进化策略的状态调整上下文指导。策略熵和提示条件注意力被用作互补信号来指导这种适应。实验表明,COEVO在固定上下文强化学习上持续改进任务性能,并产生对系统提示变化更具鲁棒性的策略。更广泛地说,我们的结果表明,外部上下文不应仅被视为大型语言模型的固定接口,而应被视为递归自我改进的自适应组件。

英文摘要

Recursive self-improvement (RSI) seeks to move large language models beyond static training pipelines toward systems that can participate in improving their own future behavior. Existing approaches largely follow two directions: updating model parameters through online learning, or improving the external context through search, reflection, and prompt optimization. Although both mechanisms can support continued improvement, they are typically studied independently. This separation overlooks an important interaction: the context shapes the experience from which a model learns, while an evolving model may interpret and utilize the same context differently over time. We therefore formulate RSI as a problem of parameter--context co-evolution, where model parameters and the learning context adapt within a shared feedback loop. We introduce COEVO, a framework that updates model parameters from on-policy experience while adapting contextual guidance according to the state of the evolving policy. Policy entropy and prompt-conditioned attention are used as complementary signals to guide this adaptation. Experiments show that COEVO consistently improves task performance over fixed-context reinforcement learning and produces policies that are more robust to changes in system prompts. More broadly, our results suggest that external context should be viewed not merely as a fixed interface to a large language model, but as an adaptive component of recursive self-improvement.

发表机构

  • Peking University(北京大学)
  • Emotional Machine(情感机器(注:此处机构名含义不明确,按直译))

机构由 AI 辅助整理,请以论文原文为准。

↑