发表机构
Dartmouth College; Carnegie Mellon University(达特茅斯学院; 卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对数学推理 SFT 中统一损失导致次优训练的问题,提出 TrimSFT 方法,按对数几率差修剪已掌握和弱支持的 token,在六个模型五个基准上平均性能最优,最高提升 26.9 点。
AI 中文摘要
监督微调(SFT)对所有目标 token 施加统一的交叉熵损失,尽管不同 token 为数学推理提供的学习信号并不相同。这种统一处理可能会过度强化已经掌握的 token,同时放大对不确定、低置信度 token 的学习压力,导致次优的训练动态。我们提出了修剪对数几率差 SFT(TrimSFT),一种简单的 token 级重加权方法,根据黄金 token 与其最强竞争者之间的对数几率差来缩放 SFT 损失。TrimSFT 将监督从两个极端修剪掉:已经掌握的 token(对数几率差大)和当前模型支持较弱的 token(对数几率差小或为负),将学习集中在两者之间的中间对数几率差区域。我们以中心为 margin m、带宽为 τ 的高斯权重实例化这一原则,无需参考模型或额外的前向传播。我们在来自 Llama、Qwen 和 DeepMath 家族的六个基础模型上,跨五个数学推理基准评估了 TrimSFT。TrimSFT 持续优于标准 SFT,在六个模型中的五个上取得了最佳平均性能,在 MATH500 上相比 SFT 最高提升了 +26.9 个点。进一步分析表明,带宽 τ 比精确的 margin 位置更重要,并且仅从一侧移除监督压力的半修剪变体产生较差的权衡。token 级对数几率差分布分析表明,TrimSFT 以比统一 SFT 或单调重加权方法更平衡的方式重塑模型置信度。这些结果表明,推理 SFT 可以通过修剪两个极端而不是统一对待所有 token 而受益。
英文摘要
Supervised fine-tuning (SFT) applies a uniform cross-entropy loss to all target tokens, even though different tokens provide unequal learning signals for mathematical reasoning. This uniform treatment can over-sharpen already mastered tokens while amplifying learning pressure on uncertain, low-confidence tokens, leading to suboptimal training dynamics. We propose Trimmed Logit-Gap SFT (TrimSFT), a simple token-level reweighting method that scales the SFT loss according to the logit gap between the gold token and its strongest competitor. TrimSFT trims supervision away from both extremes: tokens already mastered (large logit gap) and tokens weakly supported by the current model (small or negative logit gap), concentrating learning within an intermediate logit-gap region between them. We instantiate this principle with a Gaussian weight centered at margin m with bandwidth τ, requiring no reference model or additional forward pass. We evaluate TrimSFT on six base models from the Llama, Qwen, and DeepMath families across five mathematical reasoning benchmarks. TrimSFT consistently improves over standard SFT, achieving the best average performance on five out of six models, with gains of up to +26.9 points over SFT on MATH500. Further analyses show that the bandwidth τ matters more than the exact margin location, and that half-trim variants that remove supervision pressure from only one side yield inferior trade-offs. A token-level logit-gap distribution analysis suggests that TrimSFT reshapes model confidence in a more balanced way than uniform SFT or monotonic reweighting methods. These results suggest that reasoning SFT can benefit from trimming both extremes rather than treating all tokens uniformly.
CommentsAccepted to Findings of the Association for Computational Linguistics: EMNLP 2026. 14 pages. Code available at https://github.com/karpning/TrimSFT