发表机构
BrainChip Inc.(脑芯片公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究MoE模型推理时内存效率问题,提出StickyMoE方法,通过惩罚相邻token间专家切换,使专家表示和路由决策共同适应。实验表明该方法能大幅降低专家切换率,在质量-局部性方面优于事后微调。
AI 中文摘要
专家混合(MoE)模型每个token仅激活稀疏的专家子集,但连续的token经常激活不同专家,导致边缘设备上慢速存储和快速内存之间频繁进行权重交换。现有补救措施要么是系统级的(缓存启发式),要么是事后的(路由器微调),在预训练期间未改变根本原因。我们提出了StickyMoE,一种可微路由一致性损失,惩罚相邻token之间的突然专家切换,鼓励路由器在语义连贯跨度上保持相同的专家分配。StickyMoE无需架构更改,只需添加一个超参数lambda,与事后方法不同,它允许专家表示和路由决策从第一个训练步骤开始共同适应。在小规模MoE语言模型上的实验表明,StickyMoE将专家切换率降低了多达60%,困惑度下降不到4%,在质量-局部性前沿上帕累托优于事后微调。路由时间局部性在训练时最有效地得以灌输。
英文摘要
Mixture-of-Experts (MoE) models activate only a sparse subset of experts per token, yet consecutive tokens frequently activate different experts -- causing constant weight swapping between slow storage and fast memory on edge devices. Existing remedies are either system-level (caching heuristics) or post-hoc (router fine-tuning), leaving the root cause unchanged during pretraining. We propose StickyMoE, a differentiable routing consistency loss that penalises abrupt expert switches between adjacent tokens, encouraging the router to maintain the same expert assignment across semantically coherent spans. StickyMoE requires no architectural changes, adds a single hyperparameter lambda, and unlike post-hoc methods, allows expert representations and routing decisions to co-adapt from the first training step. Experiments on small-scale MoE language models show that StickyMoE reduces the expert switch rate by up to 60% with less than 4% perplexity degradation, Pareto-dominating post-hoc fine-tuning on the quality-locality frontier. Routing temporal locality is most efficiently instilled at training time.