发表机构
Rensselaer Polytechnic Institute; IBM Research(伦斯勒理工学院; IBM研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
MaskCoFT通过掩码协同微调路由器和专家,减少MoE推理中的专家获取次数,提升解码速度并保持准确率。
AI 中文摘要
混合专家(MoE)语言模型常常超出单个GPU的内存容量。专家卸载将大部分专家保存在主机内存中,并按需加载,因此解码速度取决于每个token需要获取的专家数量。缓存和预取技术只能在路由允许的范围内降低这一开销。仅对路由器进行微调可以重塑路由以复用专家,但专家保持冻结状态,无法适应新路由分配给它们的token。我们提出MaskCoFT,一种掩码协同自适应微调方法,仅使用交叉熵损失同时训练路由器和专家。在微调过程中,一个可学习的二元掩码将每层的Top-K路由限制为专家子集,专家则适应被重定向给它们的token。在推理时,学习到的掩码成为对专家重新排序的软先验,因此每个专家仍可被选择。我们针对Mixtral-8x7B模拟每层4个专家的GPU缓存,对DeepSeek-V2-Lite模拟每层12个专家。MaskCoFT相对于基础模型,将每个token的专家获取次数分别减少了23.7%和10.1%。在真实卸载系统服务中,它分别将每个输出token的时间降低了高达16.4%和5.5%。在九个基准上的平均准确率仍高于基础模型0.92和0.53个百分点。
英文摘要
Mixture-of-experts (MoE) language models often exceed the memory of a single GPU. Expert offloading keeps most experts in host memory and loads them on demand, so decoding speed depends on how many experts each token must fetch. Caching and prefetching reduce this cost only as far as the routing allows. Router-only fine-tuning can reshape the routing to reuse experts, but it keeps the experts frozen, so they cannot adapt to the tokens the new routing sends them. We propose MaskCoFT, a masked co-adaptive fine-tuning method that trains routers and experts together with the cross-entropy loss alone. During fine-tuning, a learnable binary mask restricts the Top-K routing of each layer to a subset of experts, and the experts adapt to the tokens redirected to them. At inference, the learned mask becomes a soft prior that re-ranks experts, so every expert remains selectable. We simulate a GPU cache of 4 experts per layer for Mixtral-8x7B and 12 for DeepSeek-V2-Lite. MaskCoFT cuts expert fetches per token by 23.7% and 10.1% relative to the base model. In real offloading system serving, it lowers the time per output token by up to 16.4% and 5.5%, respectively. Its average accuracy over nine benchmarks stays above the base model by 0.92 and 0.53 points.