发表机构
University of Liverpool; University of Sheffield(利物浦大学; 谢菲尔德大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出在线自加权微调(OSW-FT)方法,通过少量在线rollout估计模型成功率以调整SFT损失,在Qwen3系列模型上提升了中小规模模型在二元可验证推理任务的性能,计算性能权衡良好。
AI 中文摘要
标准监督微调(SFT)会给每一条专家演示分配相同的显式损失权重,而不考虑模型在训练查询过程中不断变化的能力。基于强化学习(RL)的方法会利用模型生成的rollout来调整更新强度,但通常需要大量采样,且在困难任务上表现不稳定。我们提出在线自加权微调(OSW-FT),这是一种简单的方法,通过在线轨迹级加权来增强SFT。对于每个查询,OSW-FT使用少量仅推理的rollout估计模型当前的成功率,并相应地重新缩放标准SFT损失。优化方向仍以专家轨迹为基准,而更新幅度则在线调整。对于二元可验证推理,我们受方差减少原则启发,在梯度层面将这种加权与SFT和RL联系起来。所得估计量对于任何有限rollout数量的精确OSW-FT替代更新都是无偏的,我们分析了其相对于相应替代目标的收敛性。在Qwen3系列(范围从0.6B到4B)上,针对多个具有挑战性的基准(例如AIME)进行评估,OSW-FT在中小规模模型上始终优于SFT。OSW-FT提供了良好的计算性能权衡,作为一种实用方法,仅需2次在线rollout即可对中小规模大语言模型(LLM)进行二元可验证推理任务的微调。
英文摘要
Standard supervised fine-tuning (SFT) assigns the same explicit loss weight to every expert demonstration, regardless of the model's changing competence over training queries. Reinforcement learning (RL) based methods adapt update strength using model-generated rollouts, but often require substantially more sampling and can be unstable on hard tasks. We propose \textbf{Online Self-Weighted Fine-Tuning (OSW-FT)}, a simple method that augments SFT with online, trajectory-level weighting. For each query, OSW-FT estimates the model's current success rate using a small number of inference-only rollouts and rescales the standard SFT loss accordingly. The optimization direction remains anchored to the expert trajectory, while the update magnitude adapts online. For binary-verifiable reasoning, we connect this weighting to SFT and RL at the gradient level, inspired by variance-reduction principles. The resulting estimator is unbiased for the exact OSW-FT surrogate update for any finite rollout count, and we analyze convergence with respect to the corresponding surrogate objective. Evaluated across Qwen3 series ranging from 0.6B to 4B on multiple challenging benchmarks (e.g., AIME), OSW-FT consistently improves over SFT on small-to-medium scale models. OSW-FT offers a favorable compute-performance trade-off as a practical approach for fine-tuning small-to-medium LLMs on binary-verifiable reasoning tasks with only \textbf{2 online rollouts}.