发表机构
University of Cambridge; Fudan University(剑桥大学; 复旦大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对代码生成中自蒸馏导致正确实现多样性收缩的问题,提出SPECTRUM方法,通过固定参考锚点的满秩近端谱调制保留正确解决方案多样性,在MBPP等基准上显著优于基线,确立了多样性保留作为递归自我改进的互补目标。
AI 中文摘要
一个从自身输出中学习的模型,其继承的不仅仅是输出的正确性:它还继承了自身产生的解决方案。我们提出了循环自蒸馏(Looped Self-Distillation),这是一种用于代码生成的自我进化框架,在该框架中,模型在固定的信息预算下,反复生成并学习自身的原始输出,而无需持续的外部评估或基于测试对生成样本进行选择。我们识别出一个重要的分离现象:正确性可以提高,而正确实现的广度却会收缩。我们引入了SPECTRUM,该方法在每一轮中从固定的参考锚点重新估计对损失敏感的关键/值几何结构,并将其转换为满秩的近端谱调制。所有生成的补全结果都用于训练一个单一的学生模型,该模型后续的推理无需任何干预。在MBPP上经过五轮实验后,SPECTRUM保留了初始模型64个样本正确抽象语法树丰富度的89.9%,而普通自蒸馏(Vanilla Self-Distillation)保留了66.4%,子空间投影对照组保留了65.5%。在匹配正确样本数量的情况下,这一优势依然存在。无需进一步训练或重新校准,所得的学生模型在HumanEval+和APPS Intro上也比普通自蒸馏实现了更高的匹配正确丰富度,展示了多样性优势的迁移。这些发现确立了正确解决方案的保留作为递归自我改进(RSI)的一个互补目标,并表明生成时的干预可以改善后续学生模型所保留的解决方案库。
英文摘要
A model that learns from its own outputs inherits more than their correctness: it inherits which solutions it produces. We formulate Looped Self-Distillation, a self-evolution framework for code generation in which a model repeatedly generates and learns from its own raw outputs, under a fixed information budget, without ongoing external assessment or test-based selection of the generated samples. We identify a consequential separation: correctness can improve while the breadth of correct implementations contracts. We introduce SPECTRUM, which re-estimates loss-sensitive key/value geometry from a fixed reference anchor at each round and converts it into full-rank proximal spectral modulation. All generated completions train a single student, whose subsequent inference requires no intervention. After five rounds of experiments on MBPP, SPECTRUM retains 89.9% of the initial model's 64-sample correct AST richness, compared with 66.4% for Vanilla self-distillation and 65.5% for a subspace-projection control. The advantage persists at matched correct-sample counts. Without further training or recalibration, the resulting student also achieves higher matched-correct richness than Vanilla SD on HumanEval+ and APPS Intro, demonstrating transfer of the diversity benefit. These findings establish correct-solution retention as a complementary objective of recursive self-improvement (RSI) and show that generation-time intervention can improve the solution repertoire retained by subsequent students.