arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

利用指令微调与合并实现推理模型适配

Leveraging Instruction Tuning and Merging for Reasoning Model Adaptation

Yu-Du Feng, Niels Mündler-Sasahara, Mark Vero, Martin Vechev

arXiv 2607.14895首次发表:更新:

AI 中文总结

研究如何利用大量未使用的监督微调数据提升推理语言模型性能,先对模型进行经典指令微调,再与原始推理模型合并,该技术在多领域提升了模型性能且成本效益高。

AI 中文摘要

推理语言模型(RLMs)在数学和编码等领域展现出了令人瞩目的性能。这些领域能可靠验证模型输出,对强化学习推动RLMs性能提升很重要。然而,在缺乏可靠验证器的领域训练RLMs仍具挑战性。同时,在可验证和不可验证领域,都存在大量带人工编写解决方案的未使用监督微调数据。本文表明这些数据可有效用于进一步提升RLMs性能。首先对RLMs进行经典指令微调,即无推理痕迹的监督微调,然后将指令微调模型与原始推理模型合并,恢复其在目标领域的推理行为。广泛评估表明该技术在可验证和难以验证的领域(包括编码和文本摘要)均提升了RLMs性能,同时保留了在其他领域的能力,且成本效益高,不到3美元就能实现改进。

英文摘要

Reasoning language models (RLMs) demonstrate impressive performance by leveraging test-time compute in the form of reasoning tokens. However, this behavior makes adapting RLMs to new domains challenging and expensive. The reason is that further training can disturb the learned behavior and degrade model performance. This makes it difficult to leverage supervised fine-tuning data with human-written solutions: although it contains high-quality annotations, it lacks reasoning tokens. In this work, we show how, despite this challenge, such data can be used efficiently for RLM adaptation. For this, we first use standard instruction tuning. Next, we leverage model merging to combine the instruction-tuned model with the original RLM, picking the merging ratio such that the resulting model's reasoning behavior on the target domain is recovered. We evaluate our method across four RLMs on coding and text summarization tasks, where it improves target-task performance by up to $11.0\%$ while preserving reasoning behavior and limiting the out-of-distribution score degradation to on average $0.7\%$. Importantly, our adaptations are efficient and economical, costing less than USD $\$10$ per model.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑