AI 中文总结
研究领域专家训练时长对大语言模型模型合并质量的影响,在多领域多模型规模上微调专家模型,保存不同训练步数的检查点并评估五种合并方法,发现方法依赖模式,表明训练时长和合并方法应联合选择。
AI 中文摘要
多任务模型合并将单独训练的专家模型组合成一个无需联合训练就能处理所有任务的单一模型。标准做法是在专家模型的最优验证损失时进行合并。我们通过系统研究领域专家的训练时长如何影响合并模型的质量来挑战这一传统做法。我们在三个模型规模(Qwen 3.5 0.8B、2B和4B)的五个领域(数学、代码、指令跟随、多语言和安全)上对专家模型进行微调,保存从最优训练步数的25%到500%的检查点,并在每个时长评估五种合并方法。我们的发现揭示了一种明显的方法依赖模式:简单平均法随着过拟合而急剧退化,而基于稀疏化的方法在远超验证最优值时达到最佳性能。我们通过偏差-方差分解分析对此进行了形式化,将其与随机森林进行类比——在随机森林中平均法受益于高方差的个体学习者。这些结果表明,训练时长和合并方法应联合而非独立选择。
英文摘要
Multi-task model merging combines separately trained expert models into a single model that handles all tasks without co-training. Standard practice merges experts at their optimal validation loss. We challenge this convention by systematically studying how training duration of domain experts affects the quality of the merged model. We fine-tune experts on five domains (Math, Code, Instruction Following, Multilingual, and Safety) across three model sizes (Qwen 3.5 0.8B, 2B, and 4B), saving checkpoints from 25% to 500% of the optimal training steps and evaluating five merging methods at each duration. Our findings reveal a striking method-dependent pattern: simple averaging degrades sharply with overfitting, while sparsification-based methods achieve their best performance well past the validation optimum. We formalize this through bias-variance decomposition analysis, drawing a parallel to random forests where averaging benefits from high-variance individual learners. These results suggest that training duration and merging method should be chosen jointly rather than independently.
CommentsAccepted to ICML 2026 workshop on weight-space symmetries