arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

BRANCH-MoE:面向大型嵌入模型的平衡感知树路由

BRANCH-MoE: Balance-Aware Tree Routing for Large Embedding Models

Gang Fu, Adel Javanmard, MohammadHossein Bateni, Vahab Mirrokni

arXiv 2610.06725首次发表:更新:

发表机构

Google Research; University of Southern California(谷歌研究院; 南加利福尼亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

BRANCH-MoE提出一种平衡感知的树形路由架构,通过指数移动平均估计分支概率,在多个数据集上保持任务质量与专家均衡利用,并支持局部共激活与通信减少。

AI 中文摘要

混合专家(MoE)层在不按比例增加每个样本计算量的情况下提高了模型容量。然而,传统的扁平路由器可能导致专家利用不均衡,并将专家视为无结构的集合,其索引不具有拓扑意义。我们提出了BRANCH-MoE,一种路由架构,将E个专家放置在深度为log2 E的二叉决策树的叶子上。在每个内部节点,分支概率以到达该节点的流量到达加权平均分数为中心。该均值使用指数移动平均估计,这促进了两个子子树的使用,而无需辅助负载均衡损失。我们表明,该移动平均估计具有显式的噪声-延迟权衡。我们证明,对于线性节点映射和对数凹到达分布,该机制可防止路由质量崩溃。我们进一步证明,在冻结路由器下,专家的执行频率控制其随机梯度收敛速率,并且当专家按树前缀分配给设备时,靠近根节点的自信决策限制了跨设备通信。我们在Criteo点击率预测、Forest Covertype、HIGGS和YearPredictionMSD上,使用E=16、top-4路由和五个随机种子,将BRANCH-MoE与Switch softmax、DeepSeek-V3动态偏置、Skywork logit归一化和确定性哈希路由进行了评估。我们的结果表明,分层路由可以保持任务质量和均衡利用,同时产生支持局部专家共激活和减少通信的拓扑。

英文摘要

Mixture-of-experts (MoE) layers increase model capacity without a proportional increase in per-example computation. However, conventional flat routers can yield imbalanced expert utilization and treat experts as an unstructured collection, whose indices carry no topological meaning. We introduce {\bf BRANCH-MoE}, a routing architecture that places \(E\) experts at the leaves of a binary decision tree of depth \(\log_2 E\). At each internal node the branching probability is centered on the arrival-weighted mean score of the traffic reaching that node. This mean is estimated using an exponential moving average, which promotes utilization of both child subtrees without an auxiliary load-balancing loss. We show that this moving-average estimate admits an explicit noise-lag trade-off. We prove that for linear node maps and log-concave arrival distributions, this mechanism prevents routing-mass collapse. We further establish that, under a frozen router, an expert's execution frequency controls its stochastic-gradient convergence rate, and that confident decisions near the root bound cross-device communication when experts are assigned to devices by tree prefix. We evaluate BRANCH-MoE against Switch softmax, DeepSeek-V3 dynamic-bias, Skywork logit-normalized, and deterministic hash routing on Criteo click-through-rate prediction, Forest Covertype, HIGGS, and YearPredictionMSD, using \(E=16\), top-\(4\) routing, and five random seeds. Our results show that hierarchical routing can preserve task quality and balanced utilization while inducing a topology that supports localized expert co-activation and reduced communication.

Comments26 pages, 7 tables, 2 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑