通过跨层共享神经专家提升Transformer中的参数利用率
Improving Parameter Utilization by Sharing Neural Experts Across Layers in Transformers
浏览论文内容
中文总结 AI 辅助
针对Transformer层间参数冗余,提出跨层共享专家的CS-MoE架构,实现更低困惑度并仅激活55%参数,为计算受限环境提供高效方案。
中文摘要 AI 辅助
基于Transformer的大型语言模型常常遭受层间参数冗余问题,即功能变换在网络深度上被冗余学习。我们提出CS-MoE,一种新颖的Transformer架构,通过跨层专家共享来解决这一低效问题。与广泛使用的混合专家(MoE)架构不同,后者在每个Transformer块末端使用层隔离的专家,CS-MoE将层独立专家与对集中式全局共享专家池的并发访问相结合。这种全局专家共享机制能够对令牌级参数激活和计算消耗(FLOPs)进行弹性控制。实验表明,CS-MoE在仅激活55%参数的情况下,实现了比同等规模密集Transformer更低的困惑度。此外,其性能随着激活专家数量的增加而单调提升,并通过在固定FLOPs预算下扩展共享池,接近消耗更多FLOPs的MoE对应模型。CS-MoE还在计算成本和模型容量之间建立了灵活的帕累托前沿,为计算受限环境提供了一种高效替代方案。
英文摘要
Transformer-based large language models often suffer from inter-layer parameter redundancy, where functional transformations are redundantly learned across network depths. We propose CS-MoE, a novel Transformer architecture featuring cross-layer expert sharing to address this inefficiency. Deviating from the widely used Mixture-of-Experts (MoE) architecture that terminates each Transformer block with layer-isolated experts, CS-MoE combines layer-independent experts with concurrent access to a centralized, globally shared expert pool. This \textit{Global Experts Sharing} mechanism enables elastic control over token-level parameter activation and computational consumption (FLOPs). Experiments demonstrate that CS-MoE achieves lower perplexity than equal-scale dense Transformers while activating only 55\% of parameters. Furthermore, its performance scales monotonically with an increased number of activated experts and approaches MoE counterparts that consume more FLOPs by expanding the shared pool with a fixed FLOPs budget. CS-MoE also establishes a flexible Pareto frontier between computational cost and model capacity, offering an efficient alternative for computation-constrained environments.
发表机构
- China Electronics Cloud Technology Co., Ltd.(中国电子云科技有限公司)
机构由 AI 辅助整理,请以论文原文为准。