发表机构
Virginia Tech(弗吉尼亚理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对推理任务的监督微调,发现数据选择的效果依赖模型容量与训练时长,高似然数据适配小型模型早期训练,低似然数据利于大型模型长期训练,提出需结合模型容量与计算预算选择数据。
AI 中文摘要
在推理监督微调中,针对同一指令的候选响应与学生模型当前分布的匹配程度存在显著差异。近期基于似然的响应选择方法表明,更接近学生分布的响应能提供更有效的监督,这催生了“高似然响应通常更适合微调”的假设。本文重新审视这一直觉,证明基于似然的数据选择的价值关键取决于模型容量和训练时长。通过在数学推理任务上开展控制实验,使用参数规模从15亿到80亿的学生模型,以及更强教师模型生成的监督信号,我们观察到清晰的“容量依赖”的“快速适配/缓慢增益”模式:高似然数据能带来更快、更稳定的早期改进,尤其对小型模型效果显著;但当训练时长足够长时,低似然数据对大型模型的益处会逐渐凸显。为解释该现象,我们分析了学习动态,发现小型模型往往无法吸收低似然监督信号,反而陷入浅层或重复行为,而大型模型在这类数据下更易向教师模型分布靠近。我们还提供了容量受限的蒸馏理论视角,阐明了数据难度、数据跨度与学生容量如何共同支配迁移效果。总体而言,我们的研究表明,推理任务中有效的数据选择应考虑模型容量和计算预算,而非仅单一偏好高似然监督。
英文摘要
In reasoning supervised fine-tuning, candidate responses for the same instruction can differ substantially in how well they match the student's current distribution. Recent likelihood-based response selection methods suggest that responses closer to the student distribution provide more effective supervision, motivating the hypothesis that high-likelihood responses may generally be preferable for fine-tuning. In this paper, we revisit this intuition and show that the value of likelihood-based data selection depends critically on model capacity and training duration. Through controlled experiments on mathematical reasoning, using students ranging from 1.5B to 8B parameters and supervision generated by stronger teacher models, we observe a clear \emph{capacity-dependent} ``{\color{SMALLCOLOR}\textbf{Fast-Fit}} / {\color{LARGECOLOR}\textbf{Slow-Gain}}'' pattern. High-likelihood data provides faster and more stable early improvements, especially for smaller models, but low-likelihood data becomes increasingly beneficial for larger models when training is allowed to continue longer. To explain this phenomenon, we analyze learning dynamics, showing that small models often fail to absorb low-likelihood supervision and instead fall into shallow or repetitive behaviors, while larger models are better able to move toward the teacher distribution under such data. We further provide a capacity-constrained theoretical view of distillation that clarifies how data difficulty, data span, and student capacity jointly govern transfer. Overall, our findings show that effective data selection for reasoning should be aware of model capacity and computing budget rather than based on a single universal preference for high-likelihood supervision.
CommentsAccepted to COLM 2026