发表机构
KAIST(韩国科学技术院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对MoE模型推理时动态路由导致的分布偏移,提出轻量级逐层分布对齐(LDA)方法,利用校准统计量修正表征,恢复性能并保持稀疏推理效率。
AI 中文摘要
专家混合(Mixture-of-Experts, MoE)架构已成为一种强大的范式,可在大型基础模型中扩展模型容量,同时保持高效推理。然而,大多数MoE模型采用固定的top-$k$专家选择策略,即使较少的专家可能就足够,也会为每个令牌分配相同的专家预算。推理时的动态top-$k$路由可以在不重新训练的情况下减少计算量,但现有方法往往忽视了因偏离训练时路由配置而导致的分布偏移。我们表明,减少激活专家数量会持续增加SMoE输出的RMS尺度和方差,导致表征不匹配,除了专家容量损失之外,还会造成下游性能下降。为解决这一可修正的组成部分,我们提出了逐层分布对齐(Layer-wise Distribution Alignment, LDA),这是一种轻量级的推理时修正方法,利用逐层校准统计量将减少路由的表征与默认配置对齐。在多个SMoE大语言模型、基准测试和路由策略上,LDA在减少路由下恢复了因分布偏移导致的性能损失的大部分,同时以可忽略的开销保持了稀疏推理效率。
英文摘要
Mixture-of-Experts (MoE) architectures have emerged as a powerful paradigm for scaling model capacity while preserving efficient inference in large foundation models. However, most MoE models use a fixed top-$k$ expert selection policy, assigning the same expert budget to every token even when fewer experts may be sufficient. Inference-time dynamic top-$k$ routing can reduce computation without retraining, but existing methods often overlook the distributional shift caused by deviating from the training-time routing configuration. We show that reducing the number of activated experts consistently increases the RMS scale and variance of SMoE outputs, inducing a representation mismatch that contributes to downstream performance degradation in addition to the loss of expert capacity. To address this correctable component, we propose Layer-wise Distribution Alignment (LDA), a lightweight inference-time correction that uses layer-wise calibration statistics to align reduced-routing representations with the default configuration. Across multiple SMoE LLMs, benchmarks, and routing strategies, LDA recovers much of the performance lost induced by the distributional shift under reduced routing while preserving sparse-inference efficiency with negligible overhead.
CommentsAccepted to Findings of EMNLP 2026