AI 中文总结
针对LLM推理受HBM容量限制问题,提出FlashAccel系统,将HBF集成到基于HBM的GPU中,通过架构支持、特殊数据布局及引入新管理层和编程模型,在100ms延迟约束下,提升了每GPU吞吐量和能源效率。
AI 中文摘要
大语言模型(LLM)推理日益受到GPU中高带宽内存(HBM)容量的限制,因为模型权重和KV缓存增长迅速。高带宽闪存(HBF)容量高于HBM且带宽相当,是容量受限的LLM推理的有前景的基础。但其访问延迟高、带宽利用率低且缺乏异构资源管理支持。我们提出FlashAccel,一个协同设计的系统,通过将HBF集成到基于HBM的GPU中,提供架构支持减轻访问延迟,通过特殊数据布局提高带宽利用率,引入存储管理层和编程模型。实验结果表明,在100ms延迟约束下,集成六个HBF堆栈到GPU中,FlashAccel的每GPU吞吐量和能源效率分别比仅使用HBM的GPU平均提高2.54倍和1.93倍。
英文摘要
Large language model (LLM) inference is increasingly limited by the capacity of High-Bandwidth Memory (HBM) in GPUs, as model weights and KV cache grow rapidly. High-Bandwidth Flash (HBF) provides higher capacity than HBM while offering comparable bandwidth, making it a promising substrate for capacity-constrained LLM inference. However, its inherently high access latency, low bandwidth utilization, and lack of support for heterogeneous resource management make it difficult to integrate HBF into GPUs for LLM inference. We present FlashAccel, a co-designed system that enables efficient LLM inference using HBF. FlashAccel integrates HBF into HBM-based GPUs, providing architectural support to mitigate access latency. It improves bandwidth utilization through specialized data layouts for both model weights and KV cache, and introduces an HBF-aware storage management layer together with a programming model to organize persistent data in HBF and coordinate heterogeneous memory resources at the system level. Experimental results demonstrate that integrating six HBF stacks into the GPU enables FlashAccel to deliver an average improvement of 2.49$\times$ and 1.93$\times$ in throughput per GPU and energy efficiency over the HBM-only GPU under a 100ms latency constraint, respectively.