arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

重新思考用于模型合并和提示学习的专家训练

Rethinking Expert Training for Model Merging with Prompt Learning

Christos Georgakilas, Aniello Panariello, Samir El Karrat Moreno, Simone Calderara, Dimosthenis Karatzas, Joost van de Weijer

arXiv 2607.24465首次发表:更新:

发表机构

Computer Vision Center; Universitat Autònoma de Barcelona; AImageLab, University of Modena and Reggio Emilia(计算机视觉中心; 巴塞罗那自治大学; 摩德纳大学和雷焦艾米利亚大学AImageLab)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究重新思考模型合并的专家训练,提出基于提示自适应的强大基线,引入双调谐专家两阶段训练策略,可减少特定任务参数更新幅度,提升合并性能,组合异构专家集时也有效。

AI 中文摘要

模型合并旨在将从共享基础模型训练的多个领域专家合并为单个多任务模型。现有方法主要关注改进合并过程,通常假设专家通过全参数微调获得。本研究重新审视模型合并的专家训练。首先表明基于提示的自适应提供了强大基线:独立学习的提示可跨任务利用,保持主干固定,避免权重合并干扰。在此基础上引入双调谐专家(DTEs),一种两阶段训练策略,先学习提示,再微调视觉编码器。这减少特定任务参数更新幅度,产生具有更高合并兼容性的专家。跨多个CLIP架构、全量微调及LoRA专家的实验表明,DTEs持续提升标准合并方法的合并性能,即使组合异构专家集也有效。

英文摘要

Model merging aims to combine multiple domain-specialized experts trained from a shared foundation model into a single multi-task model. Existing approaches largely focus on improving the merging procedure itself and typically assume experts obtained through full-parameter fine-tuning. In this work, we revisit expert training for model merging. We first show that prompt-based adaptation provides a strong baseline: independently learned prompts can be exploited across tasks while keeping the backbone fixed, avoiding the interference introduced by weight merging. Building on this observation, we introduce Dual-Tuned Experts (DTEs), a two-stage training strategy that first learns prompts and then fine-tunes the vision encoder. This reduces the magnitude of task-specific parameter updates and produces experts with higher merge compatibility. Experiments across multiple CLIP architectures, full fine-tuning, and LoRA experts show that DTEs consistently improve merged performance of standard merging approaches and remain effective even when combining heterogeneous sets of experts.

Comments14 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑