路由漂移本身不足以诊断合并MoE大语言模型中的故障
Routing Drift Alone Does Not Diagnose Failure in Merged MoE LLMs
浏览论文内容
中文总结 AI 辅助
本研究探究MoE模型合并后路由漂移是否意味着故障,提出路由分析工具包和选择性路由器修复(SRR),证明路由漂移本身不足以诊断故障,需以任务级干预效果为准。
中文摘要 AI 辅助
模型合并技术无需联合重训练即可高效整合多个专用大语言模型(LLM),但可能显著改变混合专家(MoE)模型中的专家路由。这种“路由漂移”常被解读为路由故障,由此引出一个尚不明确的基本问题:MoE合并后的路由漂移是否确实表明路由故障,以及何种证据应作为修复的正当理由?我们针对DeepSeekMoE、OLMoE和Qwen3-MoE展开研究,提出了一套路由分析工具包,用于受控反事实干预和词元级分析。通过交叉使用源模型与合并模型的路由输入和参数,我们将大多数专家重分配归因于输入偏移而非同层参数变化。然而,相对于源模型的路由差异并不能很好地预测通过恢复源路由带来的下一词元似然增益,且不同的专家选择可能产生方向相似的混合输出。因此,我们将路由故障操作性地定义为:在指定路由干预下、固定非路由参数时,任务损失可恢复。这些测试能在故意破坏路由器的情况下检测到可恢复的损失,而在所评估的合并模型中,恢复源路由并未确立可靠的任务收益。基于此,我们提出“选择性路由器修复(SRR)”作为案例研究,并发现源专家词元似然优势并不能可靠地识别有益的局部修正。综上,这些发现表明:路由漂移本身不足以作为路由故障的证据;源自源模型信息的修正必须通过其任务层面的干预效果来评判。分析工具包和SRR代码已发布。
英文摘要
Model merging efficiently combines specialized large language models (LLMs) without joint retraining, but can substantially alter expert routing in Mixture-of-Experts (MoE) models. Such \emph{routing drift} is often interpreted as routing failure, raising a fundamental question that remains unclear: \emph{does routing drift after MoE merging actually indicate routing failure, and what evidence should justify repair?} We investigate these questions across DeepSeekMoE, OLMoE, and Qwen3-MoE proposing a routing analysis toolkit for controlled counterfactual interventions and token-level analysis. By crossing source and merged router inputs and parameters, we attribute most expert reassignments to input shifts rather than parameter changes at the same layer. However, source-relative routing differences poorly predict next-token likelihood gains from source-route restoration, and different expert selections can produce directionally similar mixture outputs. We therefore operationalize routing failure as \textit{task loss recoverable under a specified routing intervention, with non-routing parameters fixed.} These tests detect recoverable loss under deliberate router corruption, whereas source-route restoration does not establish reliable task benefits in the evaluated merged models. Motivated by these, we propose \emph{Selective Router Repair (SRR)} as a case study, and find that source-specialist token-likelihood advantages do not reliably identify beneficial local corrections. Together, these findings show that \textbf{routing drift alone is insufficient evidence of routing failure}: source-informed corrections must be judged by their task-level intervention effects. The analysis toolkit and SRR code are released.