发表机构
Hong Kong Baptist University; TadReamk Limited(香港浸会大学; TadReamk有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对冻结型混合专家模型,提出分层Copula-Gumbel-Top-$K$路由方法,通过双侧依赖控制在固定每token路由律的前提下调节专家负载方差,初步验证了机制有效性但未实现任务级微调增益。
AI 中文摘要
随机Gumbel-Top-$K$路由器为混合专家(MoE)模型的每个token定义一个「路由律」:即有序专家列表与混合权重上的分布。我们探究:在严格固定每个token的完整路由律的前提下,不同token的路由选择可达到何种联合分布。我们提出一种双侧构造方法「分层Copula-Gumbel-Top-$K$(HCGT)」:在一组相关token内,可交换的高斯Copula使各专家坐标的Gumbel扰动正相关,可提升组内专家集合的一致性;在不相交的组对之间,可调的对偶构造引入可选择的负依赖量。我们证明,两种操作均能使每个token的有序Top-$K$样本、混合权重及包含概率在分布上与「路由层基于预路由logits的独立路由」完全一致,因此条件期望专家流量得以保留。我们刻画了由此产生的权衡关系:与独立路由相比,组内正耦合只会增大实际专家负载的方差;而在相同组内强度下,与平坦耦合相比,非负组间对立只会减小该方差。因此,一致性与负载分散度可通过不变约束面上的两个互补依赖旋钮进行控制。由于基础模型未被改动,这些旋钮可由一个基于冻结特征的小型控制器驱动,该控制器可通过得分函数估计器训练:冻结网络仅在正向传播中被评估,梯度被限制在控制器内。一项初步小型试点验证了该机制与训练路径,但未确立任务级微调增益。
英文摘要
A stochastic Gumbel-Top-K router defines, for every token of a mixture-of-experts (MoE) model, a routing law: a distribution over ordered expert lists and mixture weights. We ask which joint distributions over the routing choices of different tokens are reachable while every individual token's complete routing law is held exactly fixed. We give a two-sided construction, Hierarchical Copula-Gumbel-Top-K (CGA). Within a group of related tokens, an exchangeable Gaussian copula positively correlates the Gumbel perturbations at each expert coordinate, which can increase within-group expert-set coherence. Across disjoint pairs of groups, a tunable antithetic construction introduces a selectable amount of negative dependence. We prove that both operations leave each token's ordered Top-K sample, mixture weights, and inclusion probabilities identical in distribution to independent routing at a routing layer conditioned on its pre-routing logits; conditional expected expert traffic is preserved as a consequence. We characterize the resulting trade-off: positive within-group coupling can only inflate the variance of realized expert loads relative to independent routing, while nonnegative cross-group opposition can only reduce it relative to flat coupling at the same within-group strength. Coherence and load dispersion are thus controlled by two complementary dependence dials on the invariance constraint surface. Because the base model is untouched, the dials can be driven by a small controller over frozen features, trainable with a score-function estimator: the frozen network is evaluated only in the forward direction, and gradients are confined to the controller. An initial small-scale pilot validates the mechanism and the training route, but does not establish task-level fine-tuning gains.