MoEGen:用于实例自适应LoRA生成的混合专家模型
MoEGen: Mixture-of-Experts for Instance-Adaptive LoRA Generation
浏览论文内容
中文总结 AI 辅助
MoEGen是一种将MoE-based PEFT从专家选择转向专家条件参数生成的框架,通过将专家表示为可学习向量,解耦专家容量与适配器存储,在8个常识推理基准及医学、法律领域适配中均优于现有方法。
中文摘要 AI 辅助
参数高效微调(PEFT)可实现大语言模型的高效适配,但现有基于混合专家(MoE)的PEFT方法通常通过存储多个完整的LoRA专家来提升容量,导致适配器存储随专家数量线性增长,且适配范围受限于固定的专家池。本文探究基于MoE的PEFT能否生成实例特定的适配,而无需为每个专家单独存储一个LoRA模块。为解决这一问题,我们提出MoEGen,这是一种将基于MoE的PEFT从专家选择转向专家条件参数生成的适配框架。MoEGen并非将每个专家存储为完整的LoRA适配器,而是将每个专家表示为一个小型可学习向量,称为专家编码。它将每个输入路由到这些向量上,并利用它们的加权组合来条件化一个轻量级超网络,生成输入特定的低秩更新。该设计将专家容量与适配器存储解耦,同时支持实例条件适配。在8个常识推理基准上的实验表明,MoEGen在三个主干模型上均优于强大的静态及基于MoE的PEFT基线。MoEGen在医学和法律领域的联合适配中也表现出色。
英文摘要
Parameter-efficient fine-tuning (PEFT) enables efficient adaptation of large language models, but existing MoE-based PEFT methods typically improve capacity by storing multiple full LoRA experts, causing adapter storage to grow linearly with the number of experts and restricting adaptation to a fixed expert pool. We ask whether MoE-based PEFT can produce instance-specific adaptations without explicitly storing a separate LoRA module for each expert. To address this gap, we propose MoEGen, an adaptation framework that shifts MoE-based PEFT from expert selection to expert-conditioned parameter generation. Instead of storing each expert as a full LoRA adapter, MoEGen represents each expert as a small learnable vector, termed an expert code. It routes each input over these vectors and uses their weighted combination to condition a lightweight hypernetwork that generates input-specific low-rank updates. This design decouples expert capacity from adapter storage while enabling instance-conditioned adaptation. Experiments on eight commonsense reasoning benchmarks show consistent improvements over strong static and MoE-based PEFT baselines across three backbones. MoEGen also performs strongly in joint medical and legal-domain adaptation.