arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

工具调用智能体的SFT还是RL?跨数据、方法与规模的受控研究

SFT or RL for Tool-Calling Agents? A Controlled Study Across Data, Method, and Scale

Md Tahmid Rahman Laskar, Xue-Yong Fu, Shashi Bhushan TN

arXiv 2609.17848首次发表:更新:

发表机构

Dialpad Inc.(Dialpad 公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过跨数据、方法和规模的受控实验,发现LoRA监督微调在分布内最强,而强化学习在跨数据集迁移中略有优势,且LoRA优于全参数微调。

AI 中文摘要

关于训练数据、适应方法和模型规模如何共同影响语言模型智能体的工具调用性能,目前缺乏受控实验证据。我们评估了使用LoRA的监督微调(SFT)、通过组相对策略优化(GRPO)的强化学习(RL),以及SFT后接GRPO这三种方法,覆盖了从0.6B到32B参数的六个Qwen3模型,同时考察了分布内性能和跨数据集迁移。在0.6B-32B范围内,使用LoRA的SFT是分布内最强的方法,在18个实验设置中的15个中表现最佳。在跨数据集迁移上,各方法差距较小:GRPO在训练与测试数据集不同的54个设置中赢得29个,但其对SFT的平均优势不足一个百分点,而SFT->GRPO在两种比较中都很少最强。无论采用何种方法,数据集混合都能提供持续强劲的迁移效果,同时保持接近专门化分布内训练的性能。额外分析进一步证实,LoRA优于全参数微调,表明LoRA能更好地保留预训练的智能体行为。

英文摘要

Limited controlled evidence exists on how training data, adaptation method, and model scale jointly affect tool-calling performance in language-model agents. We evaluate supervised fine-tuning (SFT) with LoRA, reinforcement learning (RL) via Group Relative Policy Optimization (GRPO), and SFT followed by GRPO across six Qwen3 models from 0.6B to 32B parameters, covering both in-distribution performance and cross-dataset transfer. SFT with LoRA is the strongest in-distribution method throughout the 0.6B-32B range and best in 15 out of 18 experimental settings. On cross-dataset transfer, the methods are closer: GRPO wins 29 out of 54 settings where training and test datasets differ, but its margin over SFT averages under one point, and SFT->GRPO is rarely strongest in either comparison. Dataset mixing gives consistently strong transfer while staying close to specialized in-distribution training, regardless of method. Additional analysis further confirms that LoRA outperforms full-parameter fine-tuning, demonstrating that LoRA better preserves pretrained agentic behavior.

CommentsAccepted to the REALM Workshop at EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑