探索用于现代大语言模型推理的高带宽闪存:机遇与挑战
Exploring High-Bandwidth Flash for Modern LLM Inference: Opportunities and Challenges
浏览论文内容
中文总结 AI 辅助
本研究分析高带宽闪存(HBF)用于大语言模型(LLM)推理的机遇与挑战,发现HBF可提升LLM服务系统性能并降低GPU需求,但需维持类HBM读取带宽并提升耐用性。
中文摘要 AI 辅助
本研究探讨将高带宽闪存(HBF)用于大语言模型(LLM)推理的潜在益处与技术挑战。HBF作为缓解现代LLM服务系统内存容量瓶颈的有前景方案,已获得越来越多关注,但其益处与挑战仍未被充分研究。为填补这一空白,我们在HBF作为主要GPU内存组件处理读写操作的多种系统配置与运行场景下,全面分析了基于HBF的LLM服务系统。分析表明,尽管HBF的写入性能有限,但其可显著提升LLM服务系统的批量大小、吞吐量与灵活性,同时降低最低GPU需求;不过,要实现这些益处,关键在于维持与高带宽内存(HBM)相当的读取带宽,并需大幅提升其耐用性。
英文摘要
This work investigates the potential benefits and technical challenges of using high-bandwidth flash (HBF) for large language model (LLM) inference. HBF has gained increasing attention as a promising solution to mitigate memory-capacity bottlenecks in modern LLM-serving systems, but its benefits and challenges remain largely uninvestigated. To address this gap, we thoroughly analyze HBF-based LLM-serving systems under diverse system configurations and operating scenarios in which HBF serves as a main GPU-memory component to handle both reads and writes. Our analysis shows that, despite its limited write performance, HBF can significantly improve the batch size, throughput, and flexibility of LLM-serving systems while reducing the minimum GPU requirements, but realizing these benefits critically depends on sustaining HBM-comparable read bandwidth and requires significant endurance improvements.