arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

BOOST:并发访问主机内存与HBM以加速LLM推理

BOOST: Concurrent Access to Host Memory and HBM to Accelerate LLM Inference

Anish Saxena, Jae Hyung Ju, Hritvik Taneja, Po-An Tsai, Aamer Jaleel, Christos Kozyrakis, Moinuddin Qureshi

arXiv 2609.13592首次发表:更新:

发表机构

Georgia Tech; Nvidia Research; Stanford(佐治亚理工学院; 英伟达研究院; 斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

BOOST通过波次感知的页面分配和运行时数据管理,实现GPU对HBM和主机内存的并发按比例访问,无需修改内核即可提升LLM推理吞吐量,在Grace Hopper上平均提升31%。

AI 中文摘要

GPU内存带宽和容量限制了大型语言模型(LLM)推理的吞吐量。GPU内存系统由高带宽内存(HBM)主层级和通过CPU到GPU互连连接的主机内存次级层级组成。当前的服务系统将这两个层级分层处理:当数据适配时,它们仅从HBM提供服务;否则,在使用前将数据从主机内存预取到HBM。在这两种情况下,主机内存带宽从未得到充分利用。预取通过利用主机内存扩展了容量,但消耗了HBM带宽用于写入,减少了可用于需求加载的带宽。我们观察到,充分利用主机和HBM带宽要求每波GPU线程块并发访问两个层级,并按其带宽比例进行访问。现有的带宽比例放置策略无法提供并发性,因为它们不了解GPU波次,也不了解大型2MB GPU页面大小。本文提出了BOOST,这是第一个提供对两个GPU内存层级并发和比例访问的运行时系统,无需修改内核即可提取主机内存和HBM的组合带宽用于LLM推理。BOOST的关键见解是利用内核访问模式使页面分配和运行时数据管理具有波次感知能力。对于静态模型权重,BOOST应用基于模数的页面放置,消除了访问比率方差;对于动态供应的注意力键值(KV)对,它使空闲KV页面池具有波次感知能力。我们将BOOST集成到vLLM中,并在Grace Hopper系统上进行了评估。在等批次大小下,BOOST将每个输出令牌时间(TPOT)比仅使用HBM的服务提高了4.3%,而预取则使TPOT降低了6%。在高吞吐量服务中,BOOST平均提高了31%的吞吐量,比预取高出15%。

英文摘要

GPU memory bandwidth and capacity limit throughput in large language model (LLM) inference. The GPU memory system consists of a primary tier of high-bandwidth memory (HBM) and a secondary tier of host memory connected via CPU-to-GPU interconnect. Current serving systems treat the tiers hierarchically: they serve exclusively from HBM when data fits, and otherwise prefetch data from host memory to HBM before use. In both cases, the host memory bandwidth is never well utilized. Prefetching expands capacity by utilizing host memory, but consumes HBM bandwidth for writes, reducing the bandwidth available for demand loads. We observe that fully utilizing both host and HBM bandwidth requires each wave of GPU threadblocks to access both tiers concurrently and in proportion to their bandwidth ratio. Existing bandwidth-proportional placement strategies fail to provide concurrency because they are not aware of GPU waves, and the large 2MB GPU page size. This paper presents BOOST, the first runtime system that provides concurrent and proportional access to both GPU memory tiers, extracting the combined bandwidth of host memory and HBM for LLM inference without kernel changes. The key insight in BOOST is to use kernel access patterns to make page allocation and runtime data management wave-aware. For static model weights, BOOST applies modulo-based page placement that eliminates access-ratio variance; for dynamically provisioned attention key-value (KV) pairs, it makes the free KV page pool wave-aware. We integrate BOOST into vLLM and evaluate on a Grace Hopper system. At iso-batch size, BOOST improves Time-per-Output-Token (TPOT) by 4.3% over HBM-only serving, whereas prefetching degrades TPOT by 6%. In high-throughput serving, BOOST improves throughput by 31% on average, outperforming prefetching by 15%.

Comments15 pages, 22 Figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑