发表机构
Indian Institute of Technology Delhi; NVIDIA(印度理工学院德里分校; 英伟达)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对稀疏激活的MoE语言模型部署瓶颈问题,提出MAESTRO结构化剪枝框架,将专家激活轨迹建模为马尔可夫链以产生全局感知启发式方法,在多领域评估中性能优于基线,跨任务方差低,模型泛化更一致。
AI 中文摘要
稀疏激活的专家混合(MoE)语言模型通过每次只激活一小部分参数来实现显著的推理效率,但其完整的专家库始终驻留在内存中,造成了高昂的部署瓶颈。现有结构化剪枝方法主要针对密集变压器设计,使用局部启发式方法评估专家重要性,而忽视了MoE路由的相互依赖性质。我们引入了MAESTRO(通过基于转移的路由进行马尔可夫链近似专家稀疏化),这是一个为MoE架构设计的结构化剪枝框架,它将自回归专家激活轨迹建模为遍历马尔可夫链,其平稳分布编码跨层依赖性,产生全局感知重要性启发式方法。在包括安全、偏差和伦理在内的五个不同领域进行评估,在严格的50%压缩机制下,MAESTRO在平均性能保留率上比现有基线高出10.61%,同时跨任务方差显著降低,表明全局、路由一致的剪枝产生的模型在异构任务中更一致地泛化。
英文摘要
Sparsely-activated Mixture-of-Experts (MoE) language models achieve remarkable inference efficiency by activating only a small fraction of parameters per token, yet their full expert banks reside in memory at all times, creating a prohibitive deployment bottleneck. Existing structured pruning methods, largely designed for dense transformers, assess expert importance using locally derived heuristics that are blind to the interdependent nature of MoE routing. We introduce MAESTRO (Markov-chain Approximated Expert Sparsification via Transition-based ROuting), a structured pruning framework designed for MoE architectures that models autoregressive expert activation trajectories as Ergodic Markov chains whose stationary distributions encode cross-layer dependencies, yielding a globally aware importance heuristic. Evaluated across five diverse domains including Safety, Bias, and Ethics, MAESTRO outperforms state-of-the-art baselines by up to 10.61% in average performance retention under a strict 50% compression regime, while exhibiting substantially lower cross-task variance, indicating that global, routing-congruent pruning produces models that generalize more consistently across heterogeneous tasks.
Comments19 pages, 4 figures