arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

IntBMoE:将块级条件引入专家组合以实现全参与的混合专家模型

IntBMoE: Integrating Block-Level Conditioning into Expert Composition for Full-Participation Mixture-of-Experts

Ran Cheng, Longfei Xu, Zheng Liu, Kaikui Liu, Xiangxiang Chu

arXiv 2609.21346首次发表:更新:

发表机构

DreamX, Alibaba Group(阿里巴巴集团 DreamX)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

IntBMoE通过块级条件与稀疏执行解耦参与度、执行度和物化度,实现全参与MoE,在图像分类、语言建模和推荐任务上超越基线,并已部署于高德地图推荐系统。

AI 中文摘要

混合专家(MoE)模型扩展了容量,但现有设计无法独立设置三个量。对于单个令牌,参与度是指有多少专家对其输出贡献知识,执行度是指实际计算了多少专家(计算成本),而物化度是指必须构建和存储多少专家规模的参数集(内存成本)。稀疏路由保持了低执行度和低物化度,但缩小了参与度:对于每个令牌,只有少数专家贡献。密集输出混合恢复了完全参与,但其执行度随专家数量增长。参数合并将执行度保持在一个专家,但其物化度随路由决策数量增长。我们提出了IntBMoE,一种块条件化的MoE,通过将密集专家组合与稀疏块执行配对,解耦了这三个量。其块来自一个小型学习码本,每个条目对应一个块。在每个内部层,一个轻量级超网络将该层池中的所有专家基础合并为一个组合专家。参与度是全面的,因为每个组合专家都利用整个池。执行度保持稀疏,因为路由器将每个令牌仅发送到少数几个块。物化度是有界的,因为码本而非输入决定了块的数量。双路径残差门控(DPRG)通过乘法门控进一步耦合两个独立组合的路径。图像分类实验显示,相对于代表性的稀疏和密集MoE基线,性能持续提升。语言建模和序列推荐上的额外实验验证了其在视觉之外的泛化能力。IntBMoE已完全部署在高德地图的生成式推荐系统中,在60毫秒延迟预算下服务数亿用户,在线A/B测试中相对UVCTR提升2.4%。我们的代码可在该https URL获取。

英文摘要

Mixture-of-Experts (MoE) scales capacity, but existing designs cannot set three quantities independently. For a single token, participation is how many experts contribute knowledge to its output, execution is how many are actually computed (compute cost), and materialization is how many expert-sized parameter sets must be built and stored (memory cost). Sparse routing keeps execution and materialization low, but shrinks participation: for each token, only a few experts contribute. Dense output-mixing restores full participation, but its execution grows with the number of experts. Parameter-merging keeps execution at one expert, but its materialization grows with the number of routing decisions. We propose IntBMoE, a block-conditioned MoE that decouples all three by pairing dense expert composition with sparse block execution. Its blocks come from a small learned codebook, one per entry. At each internal layer, a lightweight hypernetwork merges all expert bases in that layer's pool into one composed expert. Participation is full, because every composed expert draws on the entire pool. Execution stays sparse, because a router sends each token to only a few blocks. Materialization is bounded, because the codebook, not the input, fixes how many blocks exist. Dual-Path Residual Gating (DPRG) further couples two independently composed paths through multiplicative gating. Experiments on image classification show consistent gains over representative sparse and dense MoE baselines. Additional experiments on language modeling and sequential recommendation validate its generalization beyond vision. IntBMoE is fully deployed in AMap's generative recommendation system, serving hundreds of millions of users under a 60ms latency budget, with a 2.4% relative UVCTR gain in online A/B testing. Our code is available at https://github.com/AMAP-ML/DreamX-Rec/.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑