当更新不再是学习:通过可学习信息增益重新思考LLM自我进化
When Updating Stops Being Learning: Rethinking LLM Self-Evolution via learnable information gain
浏览论文内容
中文总结 AI 辅助
针对LLM自我进化中的性能退化问题,提出基于可学习信息增益的整体框架及ATRI方法,通过KL散度与熵变化衡量信息增益,实现样本重加权与跨轮早停,实验验证其优越性。
中文摘要 AI 辅助
自我进化使大型语言模型(LLM)能够利用自身生成的数据进行迭代改进,但常常遭受自我进化退化问题:性能先提升,随后进入平台期,最后下降。现有方法在组件层面解决这一问题,要么针对提问者(Questioner),要么针对求解者(Solver),却忽视了自我进化是一个紧密耦合的系统。我们提出了一个基于可学习信息增益的整体框架,该增益衡量一轮相对于上一轮提供的新颖、可参数化信息的多少。理论上,该增益等于两轮数据分布之间的Kullback-Leibler散度加上它们的熵变化。实际上,通过将一个小语言模型拟合到上一轮数据,并用负对数似然对新数据进行评分来估计该增益。基于这一诊断,我们提出了ATRI(基于信息增益的自适应训练调节),它在单轮内对样本进行重新加权,并在跨轮信息增益持续较低时停止训练。在流行数据集上的实验证明了我们方法的优越性。
英文摘要
Self-evolution lets large language models (LLMs) improve iteratively using their own generated data, but often suffers from self-evolution degeneration: performance improves, plateaus, then declines. Existing methods address this issue at the component level, targeting either the Questioner or the Solver, and overlook that self-evolution is a tightly coupled system. We propose a holistic framework based on learnable information gain, which measures how much novel, parameterizable information a round provides relative to the previous round. Theoretically, this gain equals the Kullback-Leibler divergence between the two rounds' data distributions plus their entropy change. Practically, it is estimated by fitting a small language model to the previous round and scoring new data via negative log-likelihood. Based on this diagnostic, we propose ATRI (Adaptive Training Regulation via Information-gain), which reweights samples within a round and halts training across rounds when information gain remains low. Experiments on popular datasets demonstrate the superiority of our proposal.
发表机构
- Beijing University of Posts and Telecommunications(北京邮电大学)
- Beijing Academy of Artificial Intelligence(北京人工智能研究院)
- Central South University(中南大学)
- Graduate School of China Academy of Engineering Physics(中国工程物理研究院研究生院)
- The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
机构由 AI 辅助整理,请以论文原文为准。