arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ACE:跨专家适配器整合用于MoE大语言模型的参数高效微调

ACE: Adapter Consolidation across Experts for Parameter-Efficient Fine-Tuning of MoE LLMs

Ahin Lee, Sehyun Yun, Joonha Park, Taesik Gong

arXiv 2609.06072首次发表:更新:

发表机构

Ulsan National Institute of Science and Technology (UNIST)(蔚山国立科学技术院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ACE通过整合冗余专家适配器为组共享高秩LoRA模块,在相同参数预算下提升MoE模型微调精度,并实现1.31-1.48倍训练加速。

AI 中文摘要

混合专家(MoE)模型的参数高效微调(PEFT)通常为每个专家附加一个独立的低秩适配器。这种专家级设计在三个方面割裂了适配过程:容量被分散到狭窄的低秩更新中,在稀疏路由下梯度监督变得稀疏且不平衡,执行被分解为许多小型GEMM。我们发现这种专家级分离往往是不必要的,因为在微调过程中,LoRA适配器的子集会变得功能相似,揭示了专家特定适配器之间的冗余。基于这种冗余,我们提出了ACE(跨专家适配器整合),它将冗余专家分组,并在相同的PEFT预算下,用组共享的高秩LoRA模块替换其专家特定适配器。ACE进一步引入了分组适配器执行,将碎片化的专家级适配器计算整合为更少、更大的组级GEMM。在涵盖12个数据集和四个MoE骨干网络的评估中,在具有完整基线覆盖的三个骨干网络上,ACE在参数匹配的PEFT方法中取得了最高的平均准确率,同时相较于专家级LoRA提供了1.31倍至1.48倍的挂钟训练加速,且不增加峰值内存。我们的代码可在该https URL获取。

英文摘要

Parameter-efficient fine-tuning (PEFT) of mixture-of-experts (MoE) models commonly attaches a separate low-rank adapter to each expert. This expert-wise design fragments adaptation in three ways: capacity is split across narrow low-rank updates, gradient supervision becomes sparse and imbalanced under sparse routing, and execution is decomposed into many small GEMMs. We find that such expert-wise separation is often unnecessary, as subsets of LoRA adapters become functionally similar during fine-tuning, revealing redundancy among expert-specific adapters. Based on this redundancy, we propose ACE (Adapter Consolidation across Experts), which groups redundant experts and replaces their expert-specific adapters with group-shared higher-rank LoRA modules under the same PEFT budget. ACE further introduces grouped adapter execution, which consolidates fragmented expert-wise adapter computations into fewer, larger group-level GEMMs. Across evaluations covering 12 datasets and four MoE backbones, ACE achieves the highest observed mean accuracy among the parameter-matched PEFT methods on the three backbones with complete baseline coverage, while providing $1.31\times$ to $1.48\times$ wall-clock training speedup over expert-wise LoRA without increasing peak memory. Our code is available at https://github.com/UbiquitousAILab/ACE.

Comments23 pages, 13 figures. Accepted to EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑