arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

自适应互蒸馏用于大语言模型平衡多任务后训练

Adaptive Mutual Distillation for Balanced Multi-Task Post-Training of Large Language Models

Baohang Li, Xiaocheng Feng, Yichong Huang, Chengpeng Fu, Wenshuai Huo, Zekun Zhou, Zekun Yuan, Tingjia Zhang, Bing Qin

arXiv 2610.02856首次发表:更新:

发表机构

Harbin Institute of Technology; Peng Cheng Laboratory(哈尔滨工业大学; 鹏城实验室)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出自适应互蒸馏(AMD)框架,联合训练两个采用不同任务平衡策略的模型,通过短训练探针和任务级验证分数动态调整蒸馏权重,在六个基准和三个LLM骨干上优于SFT基线,合并模型平均提升2.91个百分点。

AI 中文摘要

大语言模型(LLM)的多任务后训练旨在提升模型在训练数据量不均的多个任务上的性能。现有方法主要关注在单模型训练过程中平衡任务贡献。不同的任务平衡策略可能产生具有互补优势的模型,从而为互蒸馏创造机会。然而,跨模型监督的有效性可能因任务、迁移方向和训练阶段的不同而有所差异。我们提出自适应互蒸馏(AMD),一种协作式后训练框架,该框架联合训练两个采用不同任务平衡策略的模型。AMD通过跨任务共享的短训练探针评估蒸馏权重的候选调整方案,然后利用任务级验证分数为每个任务和迁移方向选择调整方案。在六个基准和三个LLM骨干网络上,两个AMD模型均取得了比使用相同采样策略训练的监督微调(SFT)基线更高的平均基准分数。它们还优于我们实验中评估的任务平衡方法。合并两个训练后的模型可以进一步提高其平均基准分数,同时产生一个用于推理的单一模型。合并后的模型在三个骨干网络上的平均性能比多任务SFT高出2.91个百分点。

英文摘要

Multi-task post-training of large language models (LLMs) aims to improve performance across tasks with unequal amounts of training data. Existing methods focus primarily on balancing task contributions during single-model training. Different task-balancing strategies can produce models with complementary strengths, creating opportunities for mutual distillation. However, the usefulness of cross-model supervision can vary across tasks, transfer directions, and stages of training. We propose Adaptive Mutual Distillation (AMD), a collaborative post-training framework that jointly trains two models with different task-balancing strategies. AMD evaluates candidate adjustments to distillation weights through short training probes shared across tasks, then uses task-wise validation scores to select an adjustment for each task and transfer direction. Across six benchmarks and three LLM backbones, both AMD models achieve higher average benchmark scores than supervised fine-tuning (SFT) baselines trained with the same sampling strategies. They also outperform the task-balancing methods evaluated in our experiments. Merging the two trained models can further improve their average benchmark score while yielding a single model for inference. The merged models outperform multi-task SFT by an average of 2.91 points across the three backbones.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑