发表机构
Shanghai Jiao Tong University; Alibaba Group(上海交通大学; 阿里巴巴集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对MoE推理受整体专家抽象限制的问题,提出PCoMoE路径组合式执行框架,通过路径级优化等技术实现1.31倍推理加速与10%准确率提升。
AI 中文摘要
混合专家(MoE)架构通过每个token激活稀疏的专家子集,高效扩展了大语言模型(LLM)的容量。然而,现代MoE推理仍受到刚性的整体专家抽象的严重限制。现有框架将专家作为原子执行单元进行管理、调度或剪枝,这过早地固定了优化边界,且未充分挖掘细粒度的专家内部计算冗余。本研究提出PCoMoE,一种路径组合式执行框架,将MoE推理从粗粒度的专家选择转向细粒度的路径组合。PCoMoE包含专家计算的路径级公式、一种感知兼容性的分层剪枝策略以抑制低价值的路径组合,以及一个硬件友好的执行引擎,用于在严格受限的开销下利用可复用的子专家结构。实验结果表明,PCoMoE实现了高达1.31倍的端到端推理加速,同时将模型准确率提升了10%。代码可在指定URL获取。
英文摘要
Mixture-of-Experts (MoE) architectures scale Large Language Model (LLM) capacity efficiently by activating a sparse subset of experts per token. However, modern MoE inference remains heavily constrained by the rigid, whole-expert abstraction. Existing frameworks manage, schedule, or prune experts as atomic execution units, which fixes the optimization boundary too early and leaves fine-grained intra-expert computational redundancy underexplored. In this work, we present PCoMoE, a path-compositional execution framework that shifts MoE inference from coarse-grained expert selection to fine-grained path composition. PCoMoE incorporates a path-level formulation of expert computation, a compatibility-aware layer-wise pruning strategy to suppress low-value path combinations, and a hardware-friendly execution engine to exploit reusable sub-expert structures under strictly bounded overheads. Experimental results demonstrate that PCoMoE achieves up to a 1.31x end-to-end inference speedup while enhancing model accuracy by 10%. The code is available at https://github.com/gzyyy0/PCoMoE
CommentsAccepted to EMNLP 2026 Main Conference