arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.15636cs.AR

万亿参数MoE装进盒中:利用高带宽闪存解耦内存供给

Trillion-Parameter MoE in a Box: Decoupling Memory Provisioning with High-Bandwidth Flash

  • Huawei Technologies Co., Ltd.(华为技术有限公司)

机构由 AI 辅助整理,请以论文原文为准。

Pengfei Xia, Tuo Hao, Shengwei Li, Jinjing Chen, Shiru Wei, Wenjun Zou, Rui Zhang, Hui Zang

AI总结:

针对万亿参数MoE模型,通过解耦内存供给与高带宽闪存,以低带宽比率满足性能目标,并揭示带宽扩展的非比例性。

AI中文摘要:

一种面向低并发场景的万亿参数模型MoE设备,必须在一个节点上容纳数TB的权重,并以固定资源提供预填充和解码服务。结合对两个万亿参数MoE模型的算子分析、实测专家路由轨迹以及多轮智能体服务轨迹,我们探索了一个涵盖高带宽闪存(HBF)与DRAM配置、带宽暴露以及近数据计算的设计空间。我们发现状态带宽与HBF传输构成了两个大致正交的拐点,并解决了两个供给问题。问题1:一旦权重移至HBF,DRAM所需的带宽容量比是多少?在状态层256GB下限的情况下,两个模型在仅为1.4–4.0 s⁻¹的比率下即可达到1.10倍完成时间目标,这大约比HBM3e的33.3 s⁻¹低一个数量级。问题2:随着HBF内部带宽随容量扩展,面向主机的带宽是否必须按比例扩展?六个HBF封装每个向主机暴露384 GB/s,总计2.30 TB/s,比6.14 TB/s的全暴露参考值低62.5%,而更多封装在相同目标下降低了每个封装所需的带宽。

英文摘要:

An MoE appliance for trillion-parameter models at low concurrency must host terabytes of weights on one node and serve prefill and decode with fixed resources. Combining operator analysis of two trillion-parameter MoE models, a measured expert routing trace, and agentic serving traces over multiple turns, we explore a design space spanning High-Bandwidth Flash (HBF) and DRAM configurations, bandwidth exposure, and near-data compute. We find that state bandwidth and HBF transport form two largely orthogonal knees and address two provisioning questions. Q1: Once weights move to HBF, what bandwidth-to-capacity ratio does DRAM require? With a 256-GB floor for the state tier, both models meet a $1.10\times$ completion time target at ratios of only $1.4$--$4.0~\mathrm{s}^{-1}$, roughly an order of magnitude below HBM3e's $33.3~\mathrm{s}^{-1}$. Q2: As HBF internal bandwidth scales with capacity, must bandwidth to the host scale proportionally? Six HBF packages expose 384~GB/s per package to the host, 2.30~TB/s aggregate and 62.5\% below the 6.14~TB/s full exposure reference, while more packages reduce required bandwidth per package at the same target.

↑