ACE:面向基于MoE的大语言模型的自适应免校准专家跳过机制
ACE: Adaptive Calibration-Free Expert Skipping for MoE-based LLMs
AI总结:
该研究针对MoE-LLM固定top-k路由的冗余计算问题,提出免训练免校准的ACE框架,结合GSP与RCR组件实现自适应专家跳过,在多基准测试中优于现有方法,Qwen模型50%跳过率下性能提升显著。
AI中文摘要:
混合专家(MoE)架构为大语言模型(LLM)的扩展提供了高效范式,但固定的top-k路由会为每个token激活相同数量的专家槽位,造成大量冗余计算。现有的专家跳过方法通常依赖路由置信度、校准数据或额外训练,因此无法可靠估计被路由专家的实际贡献。为此,我们提出ACE,一种面向基于MoE的LLM的token自适应专家跳过的免训练、免校准且保留检查点的框架。ACE包含两个互补组件:1)全局频谱代理(GSP),其从耦合的门控、上投影和下投影以及RMSNorm缩放中估计全局变换能力;2)路由条件细化(RCR),其从中心化路由权重中构建专家特定方向原型,并沿路由偏好方向评估专家响应。推理期间,ACE将两个估计值与运行时路由门结合,仅当两种视角均判定专家槽位贡献低时才跳过该槽位,且始终保留top-1专家。所有专家统计数据均离线计算,在线阶段仅保留表查找和轻量标量运算。在三个基于MoE的LLM和八个基准上开展的大量实验表明,ACE始终优于现有静态和动态基线,在激进专家跳过下优势愈发显著。例如,在Qwen3.6-35B-A3B上采用50%跳过率时,ACE较最强竞争方法将WikiText-2困惑度降低7.96%,并将平均下游准确率提升4.15个百分点。
英文摘要:
Mixture-of-Experts (MoE) architectures provide an efficient paradigm for scaling large language models (LLMs), yet fixed top-k routing activates the same number of expert slots for every token, causing substantial redundant computation. Existing expert-skipping methods often rely on router confidence, calibration data, or additional training, and therefore cannot reliably estimate the actual contribution of routed experts. To this end, we propose ACE, a training-free, calibration-free, and checkpoint-preserving framework for token-adaptive expert skipping in MoE-based LLMs. ACE contains two complementary components: 1) Global Spectral Proxy (GSP), which estimates global transformation capacity from the coupled gate, up, and down projections together with RMSNorm scaling; and 2) Router-Conditioned Refinement (RCR), which constructs expert-specific direction prototypes from centered router weights and evaluates expert responses along routing-preferred directions. During inference, ACE combines both estimates with runtime router gates and skips an expert slot only when both views identify it as low-contribution, while always retaining the top-1 expert. All expert statistics are computed offline, leaving only table lookups and lightweight scalar operations online. Extensive experiments across three MoE-based LLMs and eight benchmarks demonstrate that ACE consistently outperforms existing static and dynamic baselines, with increasingly pronounced advantages under aggressive expert skipping. For instance, at a 50% skipping ratio on Qwen3.6-35B-A3B, ACE reduces WikiText-2 perplexity by 7.96% and improves average downstream accuracy by 4.15 percentage points over the strongest competing method.