发表机构
Central University; Innopolis University(中央大学; 因诺波利斯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究证实稀疏MoE模型各层路由状态共享通用几何结构,经对齐后可复用线性过渡,且该共享动力学为路由特有,迁移规范状态可降低模型的负对数似然损失。
AI 中文摘要
稀疏混合专家(MoE)模型在每个稀疏层使用独立参数化的路由器为每个token选择专家。已有研究表明,跨层的路由决策通常可通过早期路由信号预测,这说明各层间路由并非完全独立,但这种可预测性背后的结构仍不明确。本研究提供证据表明,各层与路由相关的状态共享一种通用几何结构,该结构被层特定坐标系所掩盖。我们分离每个路由器的控制子空间,并使用广义正交Procrustes分析将这些空间对齐为共享的规范表示。对齐后,单个线性过渡的R²值达到0.39至0.71,保留了单独拟合层特定动力学的79%至90%的预测能力,表明路由状态演变在很大程度上遵循跨层可复用的过程。随后我们探究这种共享动力学是否为路由特有,或仅反映隐藏表示的平滑演变。匹配秩比较显示,残差表示通常更易跨层预测,而路由器控制状态能更忠实地保留模型的专家选择,从而将通用跨层可预测性与路由特有信息区分开。最后,我们测试当用预测的规范状态替代原生路由状态时是否仍保持意义:迁移后的状态保留了局部路由行为,而学习到的状态演变在OLMoE上相对于简单持续减少ΔNLL达15.7%,在Phi的10个路由器范围内减少ΔNLL达6.2%。
英文摘要
Sparse mixture-of-experts (MoE) models use an independently parameterized router at each sparse layer to select experts for every token. Prior work has shown that routing decisions across depth can often be predicted from earlier routing signals, suggesting that routing is not fully independent across layers. However, the structure behind this predictability remains unclear. In this work, we provide evidence that routing-relevant states across layers share a common geometric structure that is obscured by layer-specific coordinate systems. We isolate the control subspace of each router and align these spaces into a shared canonical representation using generalized orthogonal Procrustes analysis. After alignment, a single linear transition reaches $R^2=0.39$--$0.71$ and retains 79--90\% of the predictive power of separately fitted layer-specific dynamics, indicating that much of routing-state evolution follows a reusable process across depth. We then ask whether this shared dynamics is specific to routing or simply reflects the smooth evolution of hidden representations. A matched-rank comparison shows that residual representations are often easier to predict across layers, while router-control states preserve the model's expert choices much more faithfully. This separates generic cross-layer predictability from routing-specific information. Finally, we test whether the predicted canonical states remain meaningful when used in place of native routing states. The transported states preserve local routing behavior, while learned state evolution reduces $Δ\mathrm{NLL}$ relative to simple persistence by 15.7\% on OLMoE and 6.2\% over a 10-router horizon on Phi.