发表机构
Apple(苹果公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对MoE模型结合投机解码时内存传输瓶颈问题,发现训练中采用全局负载均衡损失、共享专家、一致性损失和自回归专家选择机制可增强专家共激活,从而在保持准确率的同时将吞吐量提升21%。
AI 中文摘要
混合专家(Mixture-of-Experts, MoE)模型越来越多地与投机解码(Speculative Decoding, SD)结合使用以加速推理,但将两者结合具有挑战性。投机解码通过并行验证令牌组来提升稠密模型的推理速度。然而,对于MoE模型,投机解码的推理加速高度依赖于被验证的令牌数量。使用更多的验证令牌会导致更多专家从DRAM传输到神经处理单元(NPU),从而增加内存传输成本。由于内存传输通常是推理中的瓶颈,这会对模型运行时间产生负面影响。在本工作中,我们研究了训练期间MoE路由器的设计对MoE模型结合投机解码时速度的影响。我们发现,具有高度专家共激活的路由器能显著加快运行时间,减轻使用更多验证令牌带来的影响。基于这一观察,我们使用十亿参数规模的Transformer模型评估了各种路由器设计选择对专家共激活和运行时间的影响。我们发现,在训练期间结合全局负载均衡损失、共享专家、一致性损失和自回归专家选择机制,能显著增强专家共激活。这种增强的共激活转化为更高的整体运行吞吐量:我们的探索产生了一个模型,其吞吐量比MoE基线提高了21%,同时保持了与基线MoE相当的准确率。
英文摘要
Mixture-of-Experts (MoE) models are increasingly deployed alongside Speculative Decoding (SD) to accelerate inference, but combining the two is challenging. SD improves the inference speed of dense models by verifying groups of tokens in parallel. However, the inference speedup for SD with MoEs depends heavily on the number of tokens being verified. Using more verification tokens results in more experts being transferred from DRAM to the Neural Processing Unit (NPU), which increases the memory transfer cost. This negatively impacts model runtime, as memory transfer is typically the bottleneck in inference. In this work, we investigate the impact of MoE router design during training on the speed of MoEs with SD. We find that routers with high degrees of expert coactivation result in much faster runtimes, mitigating the impact of using more verification tokens. Motivated by this observation, we assess the impact of various router design choices on expert coactivation and runtime using billion-parameter transformer models. We find that combining a global load-balancing loss, shared experts, a consistency loss, and an autoregressive expert selection mechanism during training results in significantly stronger expert coactivation. This increased coactivation translates into higher overall runtime throughput: our exploration yields a model that improves throughput by 21% over MoE baselines, while maintaining on-par accuracy with the baseline MoE.