arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越容量:通过具有直接GPU和HBM路径的高带宽闪存实现可扩展的MoE大语言模型推理

Beyond Capacity: Scalable MoE LLM Inference via High-Bandwidth Flash with Direct GPU and HBM Paths

Seeyeon Kim, Juhyeong Jin, Joo-Young Kim

arXiv 2608.14333首次发表:更新:

AI 中文总结

该研究针对MoE大语言模型推理的HBM容量瓶颈,提出同时利用HBF到GPU的直接路径和HBF经HBM到GPU的中继路径的架构,可提升专家传递效率,实现1.94倍吞吐量与1.90倍端到端加速。

AI 中文摘要

现代混合专家(MoE)语言模型对高带宽内存(HBM)的容量和成本效率提出了越来越高的要求,因为必须在GPU附近配置快速增长的专家权重。高带宽闪存(HBF)提供了大得多的容量,但传统设计通常通过HBF驻留的专家权重传递到GPU,而未充分利用额外的直接GPU-HBF连接。我们探索一种HBF组织方式,同时利用两条独立的专家传递路径:将专家权重从HBF直接传输到GPU的直接路径,以及将专家权重从HBF通过HBM基底芯片传输到GPU的中继路径。整个专家被分配到两条路径中的一条,两条路径上的传输同时进行,在不复制专家权重或引入共享中继瓶颈的情况下提高了总专家传递带宽。提前专家确定在常规执行点之前识别即将到来的专家,使HBF读取延迟与前面的计算重叠,而不可变专家权重和可变KV缓存数据的单独管理减少了两类流量之间的干扰。我们使用具有经验测量的GPU计算延迟的事件驱动连续批处理LLM服务模拟器评估该架构。在代表性MoE工作负载中,同时利用直接GPU-HBF和HBF-HBM-GPU路径始终比仅限制在任一路径的设计提高专家传递效率。对于代表性工作负载,与将所有HBF驻留专家权重通过HBM基底芯片传递到GPU的设计相比,所提出的架构可实现1.94倍的吞吐量和1.90倍的端到端加速。

英文摘要

Modern mixture-of-experts (MoE) language models increasingly strain the capacity and cost efficiency of high-bandwidth memory (HBM), as rapidly growing expert weights must be provisioned close to GPUs. High-bandwidth flash (HBF) offers substantially greater capacity, but conventional designs typically deliver HBF-resident expert weights to the GPU through HBM, leaving an additional direct GPU-HBF connection underutilized. We explore an HBF organization that simultaneously exploits two independent expert-delivery routes: a direct path that transfers expert weights from HBF to the GPU and a relay path that transfers them from HBF through the HBM base die to the GPU. Whole experts are assigned to one of the two routes, and transfers over both routes proceed concurrently, increasing aggregate expert-delivery bandwidth without replicating expert weights or introducing a shared relay bottleneck. Early expert determination identifies upcoming experts ahead of their conventional execution point, allowing HBF read latency to overlap with preceding computation, while separate management of immutable expert weights and mutable KV-cache data reduces interference between the two traffic classes. We evaluate the architecture using an event-driven continuous-batching LLM serving simulator with empirically measured GPU compute latencies. Across representative MoE workloads, concurrently utilizing the direct GPU-HBF and HBF-HBM-GPU routes consistently improves expert-delivery efficiency over designs restricted to either route alone. For a representative workload, the proposed architecture can achieve 1.94$\times$ higher throughput and 1.90$\times$ end-to-end speedup over a design that delivers all HBF-resident expert weights to the GPU through the HBM base die.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑