发表机构
Alibaba Group; Southeast University(阿里巴巴集团; 东南大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对MoE大语言模型,提出NSFT框架,将专家细分为子专家进行参数高效微调,结合路由重要性和激活显著性选择任务相关子专家,实验证明其优于现有方法且参数更少。
AI 中文摘要
随着大语言模型(LLMs)规模的快速扩展,稠密的全参数适配变得越来越昂贵,这推动了诸如混合专家(Mixture-of-Experts,MoE)模型等稀疏和模块化架构的发展。这一转变对参数高效微调(PEFT)提出了一个关键问题:参数应以何种粒度被选择和更新?现有的PEFT方法(如LoRA)在预定义的权重矩阵上操作,而专家级稀疏微调方法则更新整个选定的专家。然而,我们观察到被激活的专家在内部是稀疏的,只有一小部分中间通道对下游任务有强烈响应,这表明专家级适配仍然过于粗糙。我们提出了NSFT(神经子专家微调),一个细粒度的PEFT框架,将MoE适配从专家细化到子专家。NSFT沿中间维度将每个专家分解为结构化的通道组,并通过结合路由重要性和专家内部激活显著性来选择与任务相关的子专家。为了优化稀疏的部分更新,NSFT进一步引入了学习率缩放和动态梯度缩放,以补偿有效更新幅度的减小。在OLMoE和Ling-mini-2.0上,跨具有挑战性的领域特定任务和通用基准的实验表明,NSFT持续优于代表性的PEFT和专家级稀疏微调基线,同时使用显著更少的可训练参数并保持有竞争力的通用能力。这些结果表明,子专家级适配是MoE大语言模型的一种更精确和高效的PEFT范式。
英文摘要
As large language models (LLMs) scale rapidly, dense full-parameter adaptation becomes increasingly expensive, motivating sparse and modular architectures such as Mixture-of-Experts (MoE) models. This shift raises a key question for parameter-efficient fine-tuning (PEFT): at what granularity should parameters be selected and updated? Existing PEFT methods such as LoRA operate on predefined weight matrices, while expert-level sparse tuning methods update entire selected experts. However, we observe that activated experts are internally sparse, with only a small fraction of intermediate channels strongly responding to downstream tasks, indicating that expert-level adaptation is still too coarse. We propose NSFT (Neural Sub-expert Fine-Tuning), a fine-grained PEFT framework that refines MoE adaptation from experts to sub-experts. NSFT decomposes each expert along the intermediate dimension into structured channel groups and selects task-relevant sub-experts by combining routing importance with intra-expert activation saliency. To optimize sparse partial updates, NSFT further introduces learning-rate scaling and dynamic gradient scaling to compensate for the reduced effective update magnitude. Experiments on OLMoE and Ling-mini-2.0 across challenging domain-specific tasks and general benchmarks show that NSFT consistently outperforms representative PEFT and expert-level sparse tuning baselines, while using substantially fewer trainable parameters and preserving competitive general capability. These results suggest that sub-expert-level adaptation is a more precise and efficient PEFT paradigm for MoE LLMs.