arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TailSFT:过滤式微调提升后训练性能

TailSFT: Filtered Fine-Tuning Improves Post-Training Performance

Sadhika Malladi, Samy Jelassi, Dylan Foster, Jordan T. Ash, Akshay Krishnamurthy

arXiv 2608.25756首次发表:更新:

发表机构

University of California San Diego; Microsoft Research NYC(加州大学圣迭戈分校; 微软研究院纽约分部)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出TailSFT过滤式微调算法,可提升模型后强化学习性能,在OLMo-3 7B上数学与代码评估pass@16最高提17%,后续GRPO的pass@1最高提4%,还引入轻量诊断方法识别适用场景。

AI 中文摘要

强化学习后训练可驱动现代AI系统的推理能力与智能体能力,但越来越多研究表明,该方法在对已有能力的基础模型进行微调时效果最佳。本文探究现有流程生成的模型是否最适合强化学习。基于先前研究强调覆盖率与pass@K是后强化学习性能的预测指标,本文对监督微调进行简单改进,提出TailSFT算法,该算法在训练过程中过滤掉已拟合的序列,从而将学习重点放在数据分布中建模不足的区域(即尾部)。本文通过受控实验与理论分析相结合的方式,对TailSFT的设计选择(尤其是具体过滤标准)进行论证与验证。在OLMo-3 7B模型上,TailSFT常可提升数学与代码评估中的pass@16性能,绝对提升幅度最高达17%,同时计算开销极小。这些更高覆盖率的检查点在后续GRPO运行中可稳定转化为最高4%的绝对pass@1提升,表明TailSFT生成的检查点是强化学习更优的初始化模型。本文还引入一种轻量诊断方法,用于识别TailSFT最可能发挥作用的场景。更广泛而言,本文的结果推动了一种有原则的、分阶段的模型开发方法,其中中间检查点的评判标准是其对后续训练的支持效果。

英文摘要

Reinforcement learning post-training drives reasoning and agentic capabilities in modern AI systems, yet a growing body of work shows that it is most effective when used to fine-tune an already capable base model. We question whether existing pipelines yield models that are most suitable for reinforcement learning. Building on prior work highlighting the role of coverage and pass@K as predictors of post-RL performance, we design a simple modification to supervised fine-tuning, TailSFT, which filters out already fit sequences during training, thereby focusing learning on under-modeled regions, or the tail, of the data distribution. We justify and validate the design choices in TailSFT, particularly the specific filtering criteria, through a combination of controlled experiments and theoretical analysis. On OLMo-3 7B, TailSFT often improves pass@16 performance on math and coding evaluations, with gains up to 17% absolute, while incurring minimal computational overhead. These higher-coverage checkpoints consistently translate to up to 4% absolute pass@1 gains in subsequent GRPO runs, demonstrating that TailSFT checkpoints serve as better initializations for RL. We further introduce a lightweight diagnostic for identifying settings where TailSFT is most likely to help. More broadly, our results motivate a principled, stage-aware approach to model development, in which intermediate checkpoints are judged by how effectively they support subsequent training.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑