arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36569cs.LGcs.AI

从检查点变化到监督微调中的选择增益

From Checkpoint Variation to Selection Gains in Supervised Fine-Tuning

  • Jilin University(吉林大学)
  • Singapore University of Technology and Design(新加坡科技设计大学)

机构由 AI 辅助整理,请以论文原文为准。

Yupeng Chang, Wenxuan Zhang, Yuan Wu

AI总结:

本研究将监督微调中的检查点选择视为有限信息决策问题,通过60条数学轨迹和跨领域复制实验证明,增加验证数据能提升选择效果,且基于生成的选择优于NLL选择,但优于最终检查点的增益尚不明确。

AI中文摘要:

检查点选择是监督微调(SFT)中的常规决策:训练产生多个检查点,但仅保留一个。然而,固定预算的比较本身并不能区分三个经验性论断:更多验证数据是否能改善检查点选择,某种选择规则是否优于验证损失选择,以及它是否优于简单保留最终检查点。因此,我们将检查点选择视为一个有限信息决策问题。在固定已完成的训练轨迹、候选检查点和独立测试项的情况下,我们改变验证预算,并分别衡量来自额外验证数据的改进、相对于负对数似然(NLL)选择的增益,以及相对于最终检查点的增益。在60条数学SFT轨迹和19种配置中,将验证预算从32个增加到305-313个样本,使生成准确率选择的独立测试准确率提高0.32个百分点(pp),检查点一致性选择提高0.29 pp,95%配置自助法置信区间分别为[0.10, 0.56]和[0.11, 0.51]。在完整验证预算下,两种基于生成的选择规则分别比匹配的NLL选择高出0.71和0.85 pp,而它们相对于最终检查点的增益仍未得到解决。在12条新训练的常识轨迹上的跨领域复制显示了相同的定性分离:将验证预算从32个增加到1,024个问题,使生成准确率和检查点一致性选择分别提高0.87和0.27 pp,而相对于最终检查点的增益再次未得到解决。综合这些结果表明,从更多验证数据中获益、优于NLL选择以及优于最终检查点是三个不同的经验性论断,需要分别的证据。

英文摘要:

Checkpoint selection is a routine decision in supervised fine-tuning (SFT): training produces multiple checkpoints, but only one is retained. Yet fixed-budget comparisons do not by themselves distinguish three empirical claims: whether more validation data improve checkpoint selection, whether a selection rule outperforms validation-loss selection, and whether it improves over simply retaining the final checkpoint. We therefore treat checkpoint selection as a finite-information decision problem. Holding completed training trajectories, candidate checkpoints, and independent test items fixed, we vary the validation budget and separately measure improvement from additional validation data, gain over negative log-likelihood (NLL) selection, and gain over the final checkpoint. Across 60 mathematical SFT trajectories and 19 configurations, increasing the validation budget from 32 to 305-313 examples raises independent-test accuracy by 0.32 percentage points (pp) for generated-accuracy selection and 0.29 pp for checkpoint agreement, with 95% configuration-bootstrap CIs of [0.10, 0.56] and [0.11, 0.50], respectively. At the full validation budget, the two generation-based rules outperform matched NLL selection by 0.71 and 0.85 pp, respectively, while their gains over the final checkpoint remain unresolved. A cross-domain replication on 12 newly trained Commonsense trajectories shows the same qualitative separation: increasing the validation budget from 32 to 1,024 questions improves generated-accuracy and checkpoint-agreement selection by 0.87 and 0.27 pp, while gains over the final checkpoint again remain unresolved. Together, these results show that benefiting from more validation data, outperforming NLL selection, and outperforming the final checkpoint are distinct empirical claims that require separate evidence.

↑