用于持续改进大语言模型的经验链(Chain-of-Experience, CoE)
Chain-of-Experience for Continual LLM Improvement
浏览论文内容
中文总结 AI 辅助
该研究提出经验链(CoE)方法,通过迭代交互让LLM从经验中持续改进,在多领域8种LLM上验证其比无反馈基线更优,结合互补反馈可进一步提升,且每令牌准确率高于现有测试时策略。
中文摘要 AI 辅助
人类会从经验中持续学习,而传统大语言模型(LLM)评估忽略了模型通过推理时交互实现改进的能力。本文研究LLM如何在测试时通过迭代经验学习,该设置被称为经验链(Chain-of-Experience, CoE),模型通过与自身或环境反馈的迭代交互积累经验轨迹,形成超越零样本推理的持续改进循环。我们用多种反馈机制实例化CoE,包括模型自反馈和环境信号(如正确性或公开代码测试通过率),并在数学、代码和知识领域用8种LLM(包括GPT-5、Gemini-2.5 Pro、Claude-4.5 Sonnet)进行评估。研究表明,利用迭代经验始终优于无反馈基线,仅自反馈即可实现显著提升,且在任务和模型上整体改进5.6%,API成本降低19%。我们进一步发现,结合互补反馈通道(如模型和正确性信号)可获得额外提升,且CoE比现有测试时策略的每令牌准确率更高。我们还观察到LLM基础能力与改进能力呈正相关,模型在弱反馈或虚假反馈下仍保持鲁棒性,不同反馈贡献不同改进方面,且大部分改进出现在迭代早期。
英文摘要
Humans continuously learn from experience, whereas conventional large language model (LLM) evaluations ignore the models' ability to improve through inference-time interaction. In this paper, we study how LLMs learn from iterative experience at test time, a setting we refer to as Chain-of-Experience (CoE), where models accumulate experiential traces through iterative interactions with self or environmental feedback to form a continual improvement loop beyond zero-shot inference. We instantiate CoE with diverse feedback mechanisms, including model self-feedback and environmental signals such as correctness or public coding test pass rates, and evaluate across math, coding, and knowledge domains using 8 LLMs, including GPT-5, Gemini-2.5 Pro, Claude-4.5 Sonnet. Our study shows that leveraging iterative experience consistently outperforms feedback-free baselines, achieving substantial gains with self feedback alone, alongside a 5.6% overall improvement and 19% lower API cost across tasks and models. We further show that combining complementary feedback channels (e.g., model and correctness signals) yields additional gains, and that CoE delivers higher accuracy per token than existing test-time strategies. We observe a positive correlation between LLM base ability and improvement capacity, and show that models remain robust under weak or spurious feedback, with different feedback contributing to distinct improvement aspects and most gains emerging early in the iterations.
发表机构
- UC Santa Cruz(加州大学圣克鲁兹分校)
- Bytedance Seed(字节跳动Seed)
机构由 AI 辅助整理,请以论文原文为准。