arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

自进化大语言模型系统何时应停止进化?

When Is Enough Enough in Self-Evolving LLM Systems?

Enoch Yin, Bin Liu, Zhengling Qi

arXiv 2610.04756首次发表:更新:

发表机构

The George Washington University; Fudan University(乔治华盛顿大学; 复旦大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对自进化LLM系统提出一种即插即用的停止与输出选择方法,通过在线序贯检验和变点估计,在保持性能的同时大幅降低计算成本。

AI 中文摘要

自进化大语言模型(LLM)系统反复提出、评估并整合对提示、技能或其他持久性工件的更新。尽管这些系统日益有效,但它们通常在预定的迭代次数或计算预算下运行,缺乏确定何时进一步进化不再有价值的原理性标准。这可能导致两个不良后果:性能饱和后仍进行不必要的计算,以及返回过度拟合或利用评估信号的后期更新的风险。这些问题促使我们研究两个基本问题:自进化系统应何时停止,以及停止后应输出什么?我们通过将第一个问题表述为在线序贯检验问题,并利用自进化LLM系统已产生的逐项配对评估结果构建一个随时有效的重启检测器来解决。我们通过将第二个问题表述为变点估计问题,并利用估计的转换点选择较早的工件进行输出来解决。所提出的程序即插即用,无需修改底层自进化算法。在两个自进化框架、三个LLM模型家族和五个基准测试中,我们的方法大幅降低了计算成本,同时保持了相当的无见测试性能。例如,在SearchQA上使用SkillOpt和DeepSeek V4 Flash时,我们的方法在第4轮停止,而非使用全部40轮预算,将令牌使用量减少了91.6%,同时实现了82.43%的无见测试准确率,而全预算运行下为82.00%。

英文摘要

Self-evolving large language model (LLM) systems repeatedly propose, evaluate, and incorporate updates to prompts, skills, or other persistent artifacts. Despite their growing effectiveness, these systems typically operate under a predetermined iteration or compute budget, without a principled criterion to determine when further evolution is no longer worthwhile. This can lead to two undesirable consequences: unnecessary computation after performance has saturated and the risk of returning late updates that overfit or exploit the evaluation signal. These issues motivate us to study two fundamental questions: when should a self-evolving system stop, and what should it output once it stops? We address the first by formulating an online sequential testing problem and constructing an anytime-valid restart detector using the per-item paired evaluation outcomes already produced by self-evolving LLM systems. We address the second by formulating a change-point estimation problem and using the estimated transition to select an earlier artifact for output. The resulting procedure is plug-and-play and requires no modification of the underlying self-evolving algorithms. Across two self-evolving frameworks, three LLM model families, and five benchmarks, our method substantially reduces computation costs while maintaining comparable unseen-test performance. For example, on SearchQA with SkillOpt and DeepSeek V4 Flash, our method stops at round 4 rather than the full budget of 40, reducing token usage by 91.6% while achieving 82.43% unseen-test accuracy versus 82.00% under the full-budget run.

Comments23 pages, 5 figures, 7 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑