arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

监督微调的后训练科学

Post-Training Science for Supervised Fine-Tuning

Charles O'Neill, Mudith Jayasekara, Harry Partridge

arXiv 2609.01244首次发表:更新:

发表机构

Baseten(巴斯滕)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过单变量搜索实验,探究监督微调中的学习率、批量大小等关键决策,在Qwen3、Llama模型及客户SFT数据集上揭示了各超参数的变化规律与模型性能的关联,为监督微调提供了可迁移的决策建议。

AI 中文摘要

每一次监督微调(SFT)运行都会面临一系列相同的决策,例如学习率、批量大小、LoRA(低秩适配)或全量微调、训练轮数、优化器以及向模型输入的数据选择。这些决策通常需要针对每一个新模型和数据集从头开始重新探索。在此,我们通过一种工具对这些决策进行测量:该工具采用一次调整一个参数的搜索方式,覆盖了Qwen3和Llama两个系列的密集模型与混合专家(MoE)模型,在四个真实场景的客户SFT数据集上,分别针对LoRA和全量微调开展实验。这些数据集构成了可控测试平台:每个任务都采用客户构建的评估方案,其训练数据由迭代式监督微调生成,该微调会不断优化模型输出直至通过评估,因此监督目标内部一致,我们所报告的任务评判标准正是数据构建时所依据的准则。我们探究了最优学习率和批量大小如何随模型规模、系列及数据变化,以及是否存在可跨场景迁移的选择规则;LoRA与全量微调的权衡关系,以及其秩和alpha参数如何决定适配器的学习能力;验证损失(或其他指标,如损失景观平坦度)是否能准确对下游质量进行排序;后训练增益是否随模型规模和数据量扩展,我们的模型阶梯通过混合专家模型扩展至2350亿参数;训练多少轮后通用指令跟随能力会下降;以及几何感知优化器是否优于AdamW。每一项建议都附带其不确定性的度量。

英文摘要

Every supervised fine-tuning run forces the same chain of decisions, such as learning rate, batch size, LoRA or full fine-tuning, how many epochs, which optimiser, and what data to feed the model. Each of these is typically rediscovered from scratch for every new model and dataset. Here we measure them under one instrument: a sweep that varies one lever at a time, and spans dense and mixture-of-experts models in two families (Qwen3 and Llama), on four real-world customer SFT datasets, for both LoRA and full fine-tuning. These datasets give a controlled testbed: each task carries an evaluation built with the customer, and its training data is produced by iterative supervised fine-tuning that refines model outputs until they pass that evaluation, so the supervised target is internally consistent and the task judge we report against is the criterion the data was built to satisfy. We ask how the optimal learning rate and batch size move with model scale, family, and data, and whether one selection rule transfers across them; what LoRA trades against full fine-tuning, and how its rank and alpha set what the adapter can learn; whether validation loss (or other metrics, such as loss landscape flatness) faithfully ranks downstream quality; whether post-training gains scale with model size and data volume, on a model ladder extended through mixtures-of-experts to 235B parameters; how many epochs to train before general instruction-following erodes; and whether a geometry-aware optimiser improves on AdamW. Each recommendation is paired with a measure of its uncertainty.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑