arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32493cs.LGcs.AI

SoFT:面向可泛化大语言模型微调的软目标

SoFT: Soft Targets for Generalizable LLM Fine-Tuning

Huihao Jing, Wenbin Hu, Shaojin Chen, Haochen Shi, Zhongwei Xie, Guijia Zhang, Yuxuan Liu, Haoyu Huang, Haoran Li, Yangqiu Song

首次发表
浏览论文内容

中文总结 AI 辅助

针对多教师多领域微调中分布内学习与分布外泛化的权衡问题,提出软目标微调(SoFT),通过设置最小目标概率和自适应正则化平衡两者,在混合推理与智能体任务上取得最佳整体性能。

中文摘要 AI 辅助

蒸馏使作为学生的语言模型能够从专家教师那里获取新能力。然而,将来自多教师、多领域示范的知识整合到单个学生模型中仍然具有挑战性。我们在此设置下研究监督式微调(SFT),其中学生必须获取多样化的能力,同时保持超出训练任务的泛化能力。我们的实验揭示了不同SFT方法在分布内学习与分布外泛化之间存在不同的权衡,这促使我们更显式地控制这一平衡。为此,我们提出了软目标微调(SoFT),以平衡从教师示范中学习与保留基础模型现有能力之间的关系。SoFT为每个示范词元设置最小目标概率,同时对基础分布进行最小的KL变化。由此产生的目标函数将示范学习与向基础模型的自适应加权正则化相结合。我们进一步使用领域特定的梯度预算来控制这一平衡,并为每条轨迹确定概率阈值。在混合领域推理和智能体任务上的实验表明,SoFT在比较方法中取得了最佳的整体性能,在分布内能力获取和分布外泛化方面均有改进。

英文摘要

Distillation enables student language models to acquire new capabilities from expert teachers. However, integrating knowledge from multi-teacher, multi-domain demonstrations into a single student remains challenging. We study supervised fine-tuning (SFT) in this setting, where students must acquire diverse capabilities while maintaining generalization beyond the training tasks. Our experiments reveal varying trade-offs between in-distribution learning and out-of-distribution generalization across SFT methods, motivating more explicit control over this balance. To this end, we propose soft-target fine-tuning (SoFT) to balance learning from teacher demonstrations with retaining the Base model's existing capabilities. SoFT sets a minimum target probability for each demonstrated token while making the smallest KL change to the Base distribution. The resulting objective couples learning from demonstrations with adaptively weighted regularization toward the Base model. We further use domain-specific gradient budgets to control this balance and determine a probability threshold for each trajectory. Experiments on mixed-domain reasoning and agentic tasks show that SoFT achieves the best overall performance among the compared methods, with improvements in both in-distribution capability acquisition and out-of-distribution generalization.

↑