发表机构
Case Western Reserve University; Kyoto University; NII LLMC(凯斯西储大学; 京都大学; 国立情报学研究所LLMC)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对循环MoE如何有效利用专家的问题,提出Foil方法:通过展平专家(减少专家层、增加每层专家数和前向传递次数)和解耦注意力(每传递独立注意力参数),在固定参数和计算量下显著降低预训练损失并改善路由,为循环MoE设计提供指导。
AI 中文摘要
循环Transformer多次重用一层模块:通过增加计算量,使固定规模的模型进一步发挥潜力,从而更充分地利用其参数;而稀疏混合专家(MoE)模型每个token仅激活众多专家中的少数几个。循环MoE融合了这两种设计理念,为MoE模型在专家利用方面带来新的潜力,但这也引发了一个问题:如何循环使用MoE?我们通过Foil方法回答了这一问题。在保持专家参数和每个token的专家计算量固定的情况下,Foil(1)展平专家,将专家层减半,每层专家数量翻倍,并增加两倍的前向传递次数,从而使每次路由决策都能从更大的专家池中选择;(2)解耦注意力,为每次传递赋予独立的注意力参数,而专家和路由器保持共享。实验表明,Foil明显优于未展平的循环基线:在200亿token时,所有Foil模型的预训练损失均低于基线;在1000亿token时,损失随展平程度的增加而单调改善,最展平的Foil在参数和计算量相同的情况下,最终损失比基线低0.012纳特,下游准确率持平或更优;解耦注意力在相同形状下也带来了更均衡且更自信的路由。我们的消融实验分析了Foil为何有效,并将发现转化为循环MoE的设计指南:循环和加宽专家层的收益相互放大,路由置信度比负载均衡更能反映健康的专家使用情况,因此稀疏循环MoE应使用更多每层专家和更多传递次数。代码和配置可在以下网址获取:https://this https URL。
英文摘要
Looped Transformers reuse one block of layers several times: by spending extra computation they push a model of fixed size further, and so use its parameters more fully; while sparse mixture-of-experts (MoE) models activate only a few of many experts for each token. Looped MoE bridges these two design philosophies and gives MoE models new potential for better expert usage, but it raises a question: how to loop a MoE? We answer it with Foil. With the expert parameters and the expert compute per token held fixed, Foil (1) flattens the experts, halving the expert layers, doubling the experts per layer and doubling the passes, so that every routing decision chooses from a larger pool, and (2) unties the attention, giving each pass its own attention parameters while the experts and routers stay shared. Experiments show that Foil clearly outperforms the unflattened looped baseline: at 20B tokens every Foil model has lower pretraining loss than the baseline; at 100B tokens the loss improves monotonically with the degree of flattening, the most flattened Foil ending 0.012 nat below the baseline at equal parameters and compute, with downstream accuracy on par or better; untying the attention also yields more balanced and more confident routing at equal shape. Our ablations analyse why Foil works and turn the findings into design guidance for looped MoE: the returns of looping and of widening the expert layers amplify each other, routing confidence tracks healthy expert use better than load balance, and a sparse looped MoE should therefore use more experts per layer and more passes. Code and configurations are available at https://github.com/SR-A-W/how-to-loop-moe.
Comments24 pages, 6 figures, 13 tables