arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

带采样的微调:SFT 比你想象的更有效

Finetuning with Sampling: SFT Learns Better Than You Think

Aayush Karan, Sitan Chen, Yilun Du

arXiv 2610.02140首次发表:更新:

发表机构

Harvard University(哈佛大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出一种MCMC采样算法,将离策略数据逐步转为在策略数据,使SFT在科学技能、数学推理等任务上媲美甚至超越RL后训练,泛化更强且遗忘更少。

AI 中文摘要

为前沿模型引入新能力一直是后训练的目标,后训练主要采用监督微调(SFT)和强化学习(RL)来实现这一目标。传统观点认为,RL 能够在新的任务上实现强泛化而不丧失现有能力,而 SFT 则容易导致泛化能力弱和灾难性遗忘。同时,SFT 可以从离策略的专家数据中学习,而 RL 必须依赖模型通过反复采样找到成功轨迹的能力。在我们的工作中,我们寻求利用在策略学习的优势,同时利用离策略数据中包含的特权信息。然而,我们不是修改学习目标以适应这些数据,而是调整数据分布以更好地适应学习者。我们引入了一种马尔可夫链蒙特卡洛(MCMC)采样算法,该算法在给定参考模型进行微调时,逐步将离策略轨迹转换为更接近在策略的轨迹。在科学技能习得、数学推理和开放式专业知识等任务中,我们的采样算法使 SFT 能够与主流后训练技术相媲美,通常在泛化上更强,遗忘上更少,优于强在策略基线。此外,由此微调得到的模型表现出强大的分布性能,并且能够超越对基础模型分布的锐化进行学习。在更高层面上,我们的方法将采样呈现为一种模型原生的算子,它塑造数据以利于可学习性,作为后训练堆栈中通用的基本原语提供更广泛的实用性。

英文摘要

Introducing new capabilities to frontier models has long been the goal of posttraining, which predominantly employs supervised finetuning (SFT) and reinforcement learning (RL) to this end. Conventional wisdom dictates that RL enables strong generalization on new tasks without losing existing capabilities, while SFT is prone to weak generalization and catastrophic forgetting. At the same time, SFT can learn from off-policy expert data, whereas RL must rely on a model's ability to find successful trajectories with repeated sampling. In our work, we seek to leverage the strength of on-policy learning while utilizing the privileged information contained in off-policy data. However, rather than modifying the learning objective to accommodate this data, we instead tailor the data distribution to better suit the learner. We introduce a Markov chain Monte Carlo (MCMC) sampling algorithm that progressively transforms off-policy traces to be more on-policy given a reference model for finetuning. Across tasks like scientific skill acquisition, mathematical reasoning, and open-ended expertise, our sampling algorithm enables SFT to rival prevailing posttraining techniques, often generalizing better and forgetting less than strong on-policy baselines. In addition, the resulting finetuned models exhibit strong distributional performance and are capable of learning beyond sharpening the base model distribution. At a higher level, our approach presents sampling as a model-native operator that shapes data for learnability, offering broader utility as a general-purpose primitive throughout the posttraining stack.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑