arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39131cs.LGcs.ARcs.DC

表征用于LLM服务的高带宽闪存

Characterizing High Bandwidth Flash for LLM Serving

  • University of California, Berkeley(加州大学伯克利分校)
  • FuriosaAI
  • ICSI(国际计算机科学研究所)
  • LBNL(劳伦斯伯克利国家实验室)

机构由 AI 辅助整理,请以论文原文为准。

Zack Yu, Chloe Wong, Coleman Hooper, Minjae Lee, Wonjun Kang, Youngjin Cho, Michael W. Mahoney, Yakun Sophia Shao, Kurt Keutzer, Amir Gholami

AI总结:

本研究针对LLM服务中的内存瓶颈,提出HBM-HBF-主机分层存储与缓冲缓存感知调度,通过轨迹驱动模拟证明HBF增强系统可显著降低完成时间并延长写入寿命。

AI中文摘要:

大语言模型(LLM)服务需要大量内存来存储模型权重和KV缓存。随着模型规模增大和上下文变长,内存容量和带宽日益成为服务性能的瓶颈。智能体工作负载通过不断增长的上下文上的重复交互加剧了这种压力,使得保留KV状态以供重用变得越来越重要。高带宽闪存(HBF)提供了一种扩展加速器内存容量以用于大语言模型(LLM)服务的方法,但其访问成本和有限的写入耐久性使其使用复杂化。我们评估了HBF在高吞吐量智能体服务中的系统设计和调度选择,以了解额外容量何时能提高服务性能和能源效率。我们引入了一个HBM-HBF-主机分层存储系统和缓冲的缓存感知调度,并使用轨迹驱动模拟来分析它们对性能、能耗和HBF写入寿命的影响。在所评估的工作负载中,最快的HBF增强系统相对于仅HBM系统将完成时间减少了36.1%-87.0%。建模的节能达到55.8%,尽管HBF在一些轻负载上增加了能耗。缓冲的缓存感知调度将评估配置中的估计HBF写入寿命从4.77年延长到14.82年。这些结果证明了协调数据放置和调度以提高服务效率同时维持实际可行的HBF写入寿命的重要性。

英文摘要:

Large language model (LLM) serving requires substantial memory to store model weights and KV caches. As models grow larger and contexts become longer, memory capacity and bandwidth increasingly become bottlenecks for serving performance. Agentic workloads compound this pressure through repeated interactions over growing contexts, making it increasingly important to retain KV state for reuse. High-bandwidth flash (HBF) offers a way to expand accelerator memory capacity for large language model (LLM) serving, but its access costs and limited write endurance complicate its use. We evaluate HBF for high-throughput agentic serving across system design and scheduling choices to understand when additional capacity improves serving performance and energy efficiency. We introduce an HBM-HBF-host hierarchical storage system and buffered cache-aware scheduling, and use trace-driven simulations to analyze their effects on performance, energy consumption, and HBF write lifetime. Across the evaluated workloads, the fastest HBF-augmented systems reduce completion time by 36.1-87.7% relative to HBM-only systems. Modeled energy savings reach 59.1%, with benefits depending on the workload and weight placement. Buffered cache-aware scheduling extends estimated HBF write lifetime from 1.21 to 14.82 years in the evaluated configuration. These results demonstrate the importance of coordinating data placement and scheduling to improve serving efficiency while sustaining a practical HBF write lifetime.

↑