发表机构
Dialpad Inc.(Dialpad公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过大规模实验比较了多种微调策略,发现多任务全参数微调是客户支持LLM的最优默认选择,并提出了顺序LoRA和模型合并等策略的实用指南。
AI 中文摘要
生产环境中的客户支持系统通常需要大语言模型(LLM)支持多种技能,例如意图分类、问答、摘要或工具使用决策。一个核心的部署问题是:这些技能应由各自的任务专用模型处理,还是应由通过多任务训练、顺序更新或模型合并训练出的单一模型处理?我们使用涵盖五个模型家族(Qwen3、Qwen3.5、Gemma-3、Llama-3.1 和 Mistral)的十三个模型(参数量从 0.6B 到 32B)以及八个客户支持数据集(包括四个公开数据集和四个专有数据集,约 74.5k 个训练样本和 8.7k 个评估样本)来研究这一问题。在固定的训练协议下,我们训练了超过 200 个检查点。实验结果表明,在我们测试的每个模型规模下,多任务全参数微调是最强的操作默认选择。专用模型在其目标任务上表现强劲,但往往在非目标任务上性能急剧下降,这使得可靠的路由机制变得重要。顺序低秩适配(LoRA)比顺序全参数微调更好地保留了早期技能,而将专用模型与其基础模型合并则能提高非目标任务上的鲁棒性,同时对于较大模型而言,在目标任务上的损失有限。最后,我们总结了在实际环境中选择微调策略的实用指南。
英文摘要
Production customer-support systems often require LLMs to support multiple skills, such as intent classification, question answering, summarization, or tool-use decisions. A central deployment question is whether these skills should be handled by separate task-specialist models or by a single model trained through multi-task training, sequential updates, or model merging. We study this question using thirteen models spanning five families (Qwen3, Qwen3.5, Gemma-3, Llama-3.1, and Mistral) from 0.6B to 32B parameters across eight customer-support datasets, spanning four public and four proprietary datasets with approximately 74.5k training and 8.7k evaluation samples. Under a fixed training protocol, we train more than 200 checkpoints. Our experiments reveal that multi-task full fine-tuning is the strongest operational default at every model size we test. Specialist models are strong on their target tasks but often degrade sharply off-task, making reliable routing important. Sequential Low-Rank Adaptation (LoRA) preserves earlier skills better than sequential full fine-tuning, while merging a specialist with its base model improves off-task robustness with limited same-task loss for larger models. We conclude with practical guidelines for selecting fine-tuning strategies in real-world settings.