arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31688cs.CL

不要重复自己:面向覆盖率的自监督微调

Don't Repeat Yourself: Self-Supervised Fine-Tuning for Coverage

Eric Fithian, Kirill Skobelev, X. Y. Han

首次发表
浏览论文内容

中文总结 AI 辅助

针对数学和编程等领域,提出DRY-SFT后训练方法,通过顺序生成不同解并独立微调,提高输出多样性和覆盖率,在多个基准上显著提升pass@100,且对模式崩溃模型效果更佳。

中文摘要 AI 辅助

在数学和编程等可验证领域中,在多次尝试中找到一种正确解可能比每次尝试的通过率更为重要。后训练可以使大语言模型的输出集中在少数几个模式上,而提高采样温度的效果有限。我们提出了“不要重复自己监督微调”(DRY-SFT),一种后训练方法,旨在提高输出多样性和覆盖率:即在多次尝试中至少有一个正确解的概率。DRY-SFT包含两个阶段。首先,对于每个问题,顺序生成K个解,向模型展示所有先前的尝试并要求生成不同的解。其次,对每个尝试独立进行微调,从上下文中移除先前的尝试。该过程不使用奖励、验证器或正确性过滤器。在HumanEval+、MBPP+和DS-1000上,DRY-SFT分别将pass@100提高了10.8、12.5和12.4个百分点,而对pass@1的代价较小。通过抽象语法树编辑距离衡量的通过解之间的结构多样性,在所有三个基准上均显著提升。DRY-SFT还解决了基础模型在同样的200次尝试中未能解决的600个问题中的244个。在九个开放权重模型中,基础模型较低的结构多样性显著预测了DRY-SFT更大的收益,表明该方法对模式崩溃更严重的模型尤其有效。

英文摘要

In verifiable domains such as math and coding, finding one correct solution among many attempts can matter more than the pass rate of each attempt. Post-training can concentrate large language model outputs around a few modes, while increasing sampling temperature has limited effectiveness. We introduce Don't Repeat Yourself Supervised Fine-Tuning (DRY-SFT), a post-training method that increases output diversity and coverage: the probability of at least one correct solution among many attempts. DRY-SFT has two stages. First, for each problem, sequentially generate K solutions, showing the model all prior attempts and asking for a different solution. Second, fine-tune on each attempt independently, removing prior attempts from the context. The process uses no reward, verifier, or correctness filter. On HumanEval+, MBPP+, and DS-1000, DRY-SFT raises pass@100 by 10.8, 12.5, and 12.4 percentage points, respectively, at a small cost to pass@1. Structural diversity, measured by abstract syntax tree edit distance among passing solutions, rises significantly on all three benchmarks. DRY-SFT also solves 244 of 600 problems that the base model did not solve in the same 200 attempts. Across nine open-weight models, lower structural diversity of the base model significantly predicts larger DRY-SFT gains, indicating that the method is especially effective on more mode-collapsed models.

补充信息

↑