发表机构
Shanghai Artificial Intelligence Laboratory; KTH Royal Institute of Technology; University of Science and Technology of China; Fudan University; The Chinese University of Hong Kong(上海人工智能实验室; 瑞典皇家理工学院; 中国科学技术大学; 复旦大学; 香港中文大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究探索MoE模型的参数高效微调方法,提出MoE$^2$-LoRA,通过双通道RCP模块耦合预训练专家专业化与任务适应性,引入全局专家池,在多MoE主干上评估,能保持通用能力并实现最优下游精度。
AI 中文摘要
专家混合(MoE)架构已在大语言模型中广泛采用,但MoE模型的参数高效微调(PEFT)仍未充分探索。现有的MoE的PEFT方法要么因使用统一适配器而忽略路由器先验,降低效率并可能遗忘,要么依赖静态专家选择,限制逐令牌能力和跨专家特征学习。本文首次尝试用MoE风格的低秩自适应微调MoE模型:名为MoE$^2$-LoRA的方法通过双通道路由条件投影(RCP)模块将预训练的专家专业化与特定任务适应性深度耦合,该模块重用基础路由器激活来为LoRA路由提供信息。还引入跨所有层共享的单个全局LoRA专家池,实现全模型自适应及均衡的专家利用。在多个不同规模和专家粒度的MoE主干上评估,MoE$^2$-LoRA在保持更强通用能力的同时始终实现了最优的下游精度。
英文摘要
Mixture-of-Experts (MoE) architectures have been widely adopted in large language models, yet parameter-efficient fine-tuning (PEFT) for MoE models remains underexplored. Existing PEFT methods for MoE either ignore router priors with uniform adapters, reducing efficiency and risking forgetting, or rely on static expert selection, limiting per-token capacity and cross-expert feature learning. In this paper, we make the first attempt to fine-tune MoE models with MoE-style low-rank adaptation: our method, entitled MoE$^2$-LoRA, deeply couples the pretrained expert specialization with task-specific adaptivity via a dual-channel Routing-Conditioned Projection (RCP) module, which reuses base router activations to inform LoRA routing. We further introduce a single global LoRA expert pool shared across all layers, enabling model-wide adaptation with emergent layer-wise affinities and balanced expert utilization. MoE$^2$-LoRA simultaneously benefits from the advantages of prior reuse, dynamic adapter routing, and model-wide knowledge sharing. Evaluated on multiple MoE backbones with varying scales and expert granularities, MoE$^2$-LoRA consistently achieves state-of-the-art downstream accuracy while retaining stronger general capabilities.
CommentsPreprint, under review