重新思考SFT的令牌加权:抑制、反转和外推已学习特征
Rethinking Token Reweighting for SFT: Suppress, Reverse, and Extrapolate Learned Features
- Shanghai AI Laboratory(上海人工智能实验室)
- University of Science and Technology of China(中国科学技术大学)
- Fudan University(复旦大学)
- KTH Royal Institute of Technology(瑞典皇家理工学院)
- The Chinese University of Hong Kong(香港中文大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出SCALE方法,通过熵引导的门控控制冻结SFT特征,实现抑制、反转和外推,提升数学推理和代码生成性能。
AI中文摘要:
监督微调(SFT)从模型认为最不可能的令牌中学习最为激进。这有助于获取新行为,但也会放大噪声或冲突的监督,并可能覆盖有用的预训练知识。通过统一的策略损失视角,我们重新审视了现有的令牌加权方法,并表明它们为演示令牌分配非负系数。因此,它们可以抑制或放大监督更新,但一旦学习了有害特征,就无法反转。此外,更大的训练权重并不等同于特征外推,因为它们改变了优化轨迹,而不是缩放固定的SFT方向。我们认为,反转和外推需要一个由固定SFT增量定义的稳定参考框架。受此启发,我们提出了SCALE(通过局部熵的选择性适应控制),一种熵引导的适应强度控制方法,它冻结预训练模型和SFT增量,仅通过最小化预测熵来学习有界的令牌和模块特定门控。这些门控根据它们与熵减少的对齐程度,抑制、反转或外推冻结的SFT特征。在Qwen2.5-Math-1.5B、Qwen2.5-Math-7B和Qwen3-4B-Base上,SCALE实现了37.84、43.60和36.57的数学推理平均值,超过了最强的相应基线,同时在通用保留基准上保持竞争力。它还在HumanEval、HumanEval+和MBPP上为所有三个模型取得了最佳的平均代码生成性能。这些结果表明,有效的SFT修正可以受益于控制已学习残差的使用方式,而不仅仅是修改它们的学习方式。
英文摘要:
Supervised fine-tuning (SFT) learns most aggressively from tokens that the model deems least likely. This helps acquire new behaviors, but also amplifies noisy or conflicting supervision and can overwrite useful pretrained knowledge. Through a unified policy-loss view, we revisit existing token-reweighting methods and show that they assign nonnegative coefficients to demonstrated tokens. Consequently, they can suppress or amplify supervised updates, but cannot reverse harmful features once learned. Moreover, larger training weights do not amount to feature extrapolation, since they change the optimization trajectory rather than scale a fixed SFT direction. We argue that reversal and extrapolation require a stable reference frame defined by a fixed SFT delta. Motivated by this, we propose SCALE (Selective Control of Adaptation via Local Entropy), an entropy-guided adaptation-strength-control method that freezes the pretrained model and the SFT delta and learns bounded token- and module-specific gates by minimizing predictive entropy alone. These gates suppress, reverse, or extrapolate frozen SFT features according to their alignment with entropy reduction. Across Qwen2.5-Math-1.5B, Qwen2.5-Math-7B, and Qwen3-4B-Base, SCALE achieves mathematical-reasoning averages of 37.84, 43.60, and 36.57, exceeding the strongest corresponding baselines while remaining competitive on general-retention benchmarks. It also attains the best average code-generation performance across HumanEval, HumanEval+, and MBPP for all three models. These results suggest that effective SFT correction can benefit from controlling how already learned residuals are used, rather than only modifying how they are learned.