路由器先验偏差:在MoE后训练中保留基础路由结构
Router Prior Bias: Preserving Base Routing Structure in MoE Post-Training
浏览论文内容
中文总结 AI 辅助
提出路由器先验偏差(RPB),在后训练中软性保持基础路由结构,优于负载均衡损失,提升MoE模型性能。
中文摘要 AI 辅助
混合专家(MoE)预训练依赖于辅助负载均衡损失(LBL)来推动每个专家的利用率趋于均匀。后训练则面临不同的情况:基础路由器已经编码了非均匀的专家共激活结构,而重新施加的均匀性目标会将其抹平。我们表明,下游性能反而依赖于软性地保持这种继承的路由结构,我们将这一原则称为软路由器锚定,并将其实例化为路由器先验偏差(RPB),这是一种训练时偏差,将路由器logits拉向从冻结的基础路由器读取的先验,同时保持路由器本身可训练。在Moonlight-16B-A3B的数学后训练中,RPB在领域内准确率达到45.77,而重新应用LBL时为31.91,无锚定微调时为29.44,并且比两者保留了更多的领域外能力。与LBL的排序在第二个模型家族(Qwen3-30B-A3B-Base)上重现,且相对于LBL的优势可在独立来源的语料库上得到验证。基于路由器权重、logits或输出分布定义的锚定表现相当,没有一致的排序,这表明效果在于约束的软性而非RPB提供的特定先验。在专家共激活图中保留的社区结构与这些收益同步,只要基础路由器的非均匀性足以形成社区,然而将相同的先验作为硬分配强制执行会保留该结构但性能急剧下降。因此,社区结构是软锚定的足迹而非其来源,实际教训是后训练期间应软性地保持继承的路由,因为将其向均匀性抹平或绝对强制执行都会带来下游成本。我们的代码将在此https URL发布。
英文摘要
Mixture-of-Experts (MoE) pretraining relies on an auxiliary load-balancing loss (LBL) to drive per-expert utilization toward uniformity. Post-training inherits a different situation: the base router already encodes non-uniform expert co-activation structure, which a re-imposed uniformity objective flattens away. We show that downstream performance depends instead on holding this inherited routing softly, a principle we term soft router anchoring, and instantiate it as Router Prior Bias (RPB), a training-time bias that pulls the router logits toward a prior read off the frozen base router while leaving the router itself trainable. On math post-training of Moonlight-16B-A3B, RPB attains 45.77 in-domain accuracy against 31.91 under re-applied LBL and 29.44 under unanchored fine-tuning, and retains more out-of-domain capability than either. The ordering against LBL reproduces on a second model family (Qwen3-30B-A3B-Base), and the advantage over LBL is resolvable on an independently sourced corpus. Anchors defined on the router weights, on its logits, or on its output distribution perform comparably with no consistent ordering, which places the effect in the softness of the constraint rather than in the particular prior RPB supplies. Retained community structure in the expert co-activation graph tracks these gains wherever the base router is non-uniform enough for communities to form, yet enforcing the same prior as a hard assignment preserves that structure while performance falls sharply. Community structure is therefore a footprint of soft anchoring rather than its source, and the practical lesson is that inherited routing should be held softly during post-training, since both flattening it toward uniformity and enforcing it absolutely carry a downstream cost. Our code will be released at https://github.com/naver-ai/rpb.
发表机构
- NAVER Applied AI Group(NAVER应用人工智能集团)
- Healthcare AI Research Institute (HARI), Seoul National University Hospital(首尔大学医院医疗人工智能研究所(HARI))
- NAVER AI Lab(NAVER人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。