arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

以监督微调为上下文缓解监督微调中的遗忘问题

SFT-as-Context Mitigates Forgetting in Supervised Fine-Tuning

Kenan Tang, Andong Hua, Chengxuan Qian, Saket Tiwari, Yao Qin

arXiv 2610.11132首次发表:更新:

发表机构

University of California, Santa Barbara(加州大学圣巴巴拉分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出SFT-as-context方法,通过将SFT模型响应作为上下文,在保留母模型通用能力的同时获取微调能力,可解决需两种能力的查询,还扩展至提升闭源LLM性能,且有理论与注意力机制支撑。

AI 中文摘要

监督微调(SFT)为大语言模型(LLM)赋予了特定领域的能力,但往往会以遗忘其母模型(即微调前的预训练模型)的通用能力为代价。这种权衡对那些同时需要特定能力和通用能力的查询场景限制尤为明显。我们提出了SFT-as-context,这是一种无需额外训练的方法,其中母模型会将SFT模型的响应作为上下文来回答查询。该方法使母模型能够通过上下文学习从SFT响应中获取微调后的能力,同时保留自身的通用能力。在19组母模型-SFT模型对和11个基准测试中,SFT-as-context在微调能力上与SFT模型的表现接近,在AIME 2024和LiveCodeBench上的差距仅为2.2和2.1个百分点,在NutriBench-English上的宏观平均绝对误差(macro MAE)差距为2.0;而在通用能力上,其与母模型的平均差距在2.2个百分点以内。值得注意的是,该方法能够解决同时需要微调能力和通用能力的查询,即便母模型和SFT模型单独都无法成功。该方法还可扩展至母模型-SFT模型对之外的场景:小型开源SFT模型的响应能够提升强大的闭源LLM的表现,性能优于两者单独使用的效果。此外,我们使用贝叶斯框架推导了理论保证,界定了SFT-as-context在微调能力上相对于SFT模型、在通用能力上相对于母模型的误差范围。我们还对注意力权重进行了可视化,发现母模型会更多地关注有用的SFT响应,而较少关注无关的响应,这表明选择性注意力有助于母模型通过上下文学习利用SFT响应。

英文摘要

Supervised fine-tuning (SFT) equips large language models (LLMs) with specialized capabilities, but often comes at the cost of forgetting the general capabilities of their parent models (i.e., the pretrained models before fine-tuning). This trade-off is especially limiting for queries that require both specialized and general capabilities. We introduce SFT-as-context, a training-free method in which the parent model uses the SFT model's response as context to answer the query. This allows the parent model to acquire fine-tuned capabilities from the SFT response through in-context learning while preserving its own general capabilities. Across 19 parent-SFT model pairs and 11 benchmarks, SFT-as-context remains close to the SFT models on fine-tuned capabilities, with gaps of only 2.2 and 2.1 percentage points on AIME 2024 and LiveCodeBench and 2.0 macro MAE on NutriBench-English, while staying within 2.2 percentage points of the parent models on general capabilities on average. Remarkably, it can solve queries requiring both fine-tuned and general capabilities, even when neither the parent nor SFT model succeeds alone. This approach also extends beyond parent-SFT pairs: responses from a small open-source SFT model can improve a strong closed-source LLM, outperforming either model alone. Furthermore, we use a Bayesian framework to derive theoretical guarantees that bound the error of SFT-as-context relative to the SFT model on fine-tuned capabilities and to the parent model on general capabilities. In addition, we visualize the attention weights and find that the parent model attends more to useful SFT responses and less to irrelevant ones, suggesting that selective attention helps the parent model use the SFT response through in-context learning.

Comments37 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑