arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大语言模型中混合专家架构的演进:路由、拓扑、负载均衡与专家并行

The Evolution of Mixture-of-Experts Architectures in Large Language Models: Routing, Topology, Load Balancing, and Expert Parallelism

Jiguo Li

arXiv 2608.08650首次发表:更新:

AI 中文总结

本技术综述梳理了大语言模型混合专家架构的演进,通过五个维度和四个控制平面分析关键技术,揭示其从激活稀疏参数到解耦路由、预算与执行的趋势。

AI 中文摘要

混合专家(Mixture-of-Experts,MoE)模型在保持每个token激活计算量受限的同时提升了参数容量,但其架构演进无法仅通过按时间顺序排列的模型发布来解释。本技术综述综合了主要论文、官方技术报告及现有综述,沿五个耦合维度组织现代混合专家系统:专家粒度、专家拓扑、路由自由度、负载均衡范围及执行结构。我们将八个架构里程碑描述为包含六条主线发展和两条正交分支的依赖图,而非八个连续代际。随后,我们通过四个控制平面分析各个系统:专家拓扑、路由、均衡与专家并行。这些平面明确了专家的存在形式、处理每个token的专家、如何控制聚合负载以及如何将选定计算映射到物理设备。该框架将Top-k路由、共享专家、细粒度专家、动态专家组合等算法选择,与token调度、设备放置、全对通信、通信-计算重叠等系统问题关联起来。我们以等预算预训练实验、质量与系统指标及开放研究问题作为结论,主要趋势是从仅激活更多稀疏参数,转向解耦语义路由、计算预算与物理执行。

英文摘要

Mixture-of-Experts models increase parameter capacity while keeping the computation activated by each token bounded, but their architectural evolution cannot be explained by a chronological list of model releases alone. This technical survey synthesizes primary papers, official technical reports, and prior surveys to organize modern Mixture-of-Experts systems along five coupled dimensions: expert granularity, expert topology, routing freedom, the scope of load balancing, and execution structure. We describe eight architectural milestones as a dependency graph with six mainline developments and two orthogonal branches, rather than as eight successive generations. We then analyze individual systems through four control planes: Expert Topology, Routing, Balance, and Expert Parallelism. These planes specify which experts exist, which experts process each token, how aggregate load is controlled, and how selected computation is mapped onto physical devices. The framework connects algorithmic choices such as Top-k routing, shared experts, fine-grained experts, and dynamic expert composition with systems concerns including token dispatch, device placement, all-to-all communication, and communication-computation overlap. We conclude with equal-budget pretraining experiments, quality and systems metrics, and open research questions. The main trend is a shift from merely activating more sparse parameters toward decoupling semantic routing, computational budgets, and physical execution.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑