arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.17940cs.LG

超越前一层:稀疏MoE路由中的残差预测结构

Beyond the Previous Layer: Residual Predictive Structure in Sparse MoE Routing

Hao Li, Yasuyuki Tahara, Yuichi Sei

首次发表
浏览论文内容

中文总结 AI 辅助

本研究探究稀疏MoE路由中,更早的专家选择是否比仅用前一次选择更能预测下一路由器,实验表明扩展历史可显著提升预测精度,揭示了超越相邻层的残差预测结构。

中文摘要 AI 辅助

稀疏混合专家模型通过一系列专家选择序列来路由每个词元。我们探究的问题是:在预测下一个路由器时,紧邻的前一次选择是否足以概括这一轨迹。利用冻结的OLMoE和JetMoE模型,我们在保留最近一次选择作为共同基线的条件下,测量了更早专家选择所带来的保留数据预测增益。在OLMoE中,将历史从一层扩展到十一层,路由器logit的$R^2$从0.59879提升至0.66544。一项预注册的JetMoE复制实验在两个目标深度上分别获得了0.14275和0.20528的四层增益,配对bootstrap区间均高于零。这些增益在非线性解码下依然存在:向小型多层感知机添加历史信息使$R^2$提升了0.17137和0.21861,而仅对最近状态进行非线性解码相对于线性探针仅增加了0.00139和0.00936。参数匹配的对照组保持了这一优势,交叉拟合的历史残差预测目标残差的$R^2$分别为0.20549和0.23556。这些发现揭示了专家选择轨迹中超越相邻层持续性的残差预测结构。

英文摘要

Sparse mixture-of-experts models route each token through a sequence of expert selections. We ask whether the immediately preceding selection adequately summarizes this trajectory for predicting the next router. Using frozen OLMoE and JetMoE models, we measure the held-out predictive gain from earlier expert selections while retaining the most recent selection as a common baseline. In OLMoE, extending the history from one to eleven layers raises router-logit $R^2$ from 0.59879 to 0.66544. A preregistered JetMoE replication yields four-layer gains of 0.14275 and 0.20528 at two target depths, with paired bootstrap intervals above zero. These gains survive nonlinear decoding: adding history to a small multilayer perceptron improves $R^2$ by 0.17137 and 0.21861, whereas nonlinear decoding of the recent state alone adds 0.00139 and 0.00936 over a linear probe. Parameter-matched controls preserve the advantage, and cross-fitted history residuals predict target residuals with $R^2$ of 0.20549 and 0.23556. These findings identify residual predictive structure in expert-selection trajectories beyond adjacent-layer persistence.

发表机构

  • The University of Electro-Communications(电气通信大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑