发表机构
Fudan University; Shanghai Innovation Institute(复旦大学; 上海创新研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出路由有效秩作为无标签诊断指标,揭示MoE推理群体在推理中呈现集中-分化-再集中的规律,并分解共模与残差贡献,定位推理努力对分化时机的影响。
AI 中文摘要
测试时扩展产生推理回滚群体,但目前尚无标准的无标签方法描述其内部计算在推理展开过程中如何重组。我们引入了路由有效秩d_eff,即基于MoE专家路由相似性构建的跨回滚图的熵有效维度。在十种MoE配置和五个数学/科学基准上,d_eff展现出可复现的低-高-低轨迹,在3,105个模型-问题群体中,98.5%的案例存在显著的内极大值:路由相似性在早期集中,在中等预算时达到最大分化,随后再次集中,且该最大值的时间随架构和推理努力程度系统性变化。一个精确的分解将群体范围的共模质量与残差谱维度分离:共模重新分配约占轨迹的三分之二,而残差谱贡献约四分之一,并在共模之外保留显著变化。该分解进一步定位行为:在非一致群体中,共模集中的增加强烈预测相同答案的可恢复性,更高的推理努力将最大值延迟2.59个八度(令牌预算翻倍),并在所有四种测试架构中一致扩展高秩周期,将努力效应定位于时间和持续时间而非峰值幅度。正确性比较将结构监控与答案选择分离,将路由有效秩定位为群体组织的可分解、无标签诊断工具——一个关于MoE推理群体在推理时间内如何分化并重新集中的原则性谱视角。
英文摘要
Test-time scaling produces cohorts of reasoning rollouts, yet there is no standard label-free account of how their internal computation reorganizes as inference unfolds. We introduce routing effective rank deff, the entropy-effective dimensionality of a cross-rollout graph built from MoE expert-routing similarity. Across ten MoE configurations and five math/science benchmarks, deff exhibits a reproducible low-high-low trajectory, with a prominent interior maximum in 98.5% of 3,105 model-question cohorts: routing similarity is concentrated early, maximally differentiated at intermediate budgets, and reconcentrated later, and the timing of this maximum varies systematically with architecture and reasoning effort. An exact decomposition separates cohort-wide common-mode mass from residual spectral dimensionality: common-mode reallocation accounts for about two thirds of the trajectory, while the residual spectrum contributes about one quarter and retains substantial variation beyond the common mode. The decomposition further localizes behavior: among non-unanimous cohorts, increases in common-mode concentration strongly predict same-answer recoverability, and higher reasoning effort delays the maximum by 2.59 octaves (doublings of the token budget) and consistently expands the high-rank period across all four tested architectures, locating the effort effect in timing and duration rather than peak amplitude. Correctness comparisons separate structural monitoring from answer selection, positioning routing effective rank as a decomposable, label-free diagnostic of cohort organization - a principled spectral lens on how MoE reasoning cohorts differentiate and reconcentrate over inference time.