arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.25782cs.AR

面向智能体LLM服务的HBM与高带宽闪存热-冷分层

Hot-Cold Tiering of HBM and High Bandwidth Flash for Agentic LLM Serving

Jongjin Baek, Won Ji, Seungjae Yoo, Joo-Young Kim

首次发表
浏览论文内容

中文总结 AI 辅助

针对智能体LLM服务中KV状态管理问题,提出利用HBM与高带宽闪存构建热-冷KV层次结构,将热集置于HBM、冷池置于HBF,实现低延迟、高并发并降低功耗。

中文摘要 AI 辅助

大语言模型(LLM)服务日益具有智能体特性,多轮会话在动作之间处于空闲状态,但必须保留其完整上下文。有限的GPU内存容量迫使不活跃的KV状态被驱逐,因此恢复会话要么导致昂贵的重计算,要么导致缓慢的互连传输。为了解决这个问题,高带宽闪存(HBF)——一种封装内的3D-NAND存储器,其容量比高带宽内存(HBM)高出数个数量级,且读取带宽相当——已成为一个强有力的候选方案。然而,其高读取能耗和有限的写入耐久性使得它无法服务于所有KV流量。幸运的是,我们的分析表明,智能体KV状态表现出不同的访问模式:一个小的热集合在每个解码步骤中被读取,而一个大的冷池仅在暂停的会话恢复时被读取。利用这一点,我们将热集合放置在HBM中,将冷池放置在HBF中,在GPU内存层级内形成热-冷KV层次结构。在采用Qwen3-Coder-30B-A3B的智能体工作负载上,我们的设计实现了14毫秒的令牌间隔时间(TBT),并在预填充之上仅增加约0.1毫秒的恢复延迟,同时每个GPU承载的并发会话数增加了24倍。通过将HBM限制在热集合,与从闪存服务所有KV相比,我们的设计还将每个8-GPU节点的读取功耗降低了7.6千瓦,确立了HBF作为HBM的冷层补充而非替代品的地位。

英文摘要

Large language model (LLM) serving is increasingly agentic, with multi-turn sessions that idle between actions yet must retain their full context. Limited GPU memory capacity forces inactive KV states to be evicted, so resuming a session incurs either costly recomputation or slow interconnect transfers. To address this, high bandwidth flash (HBF)-an on-package 3D-NAND memory offering orders-of-magnitude greater capacity than high bandwidth memory (HBM) at comparable read bandwidth-has emerged as a strong candidate. However, its high read energy and limited write endurance make it impractical to serve all KV traffic. Fortunately, our analysis shows that agentic KV states exhibit distinct access patterns: a small hot set is read for every decoding step, while a large cold pool is read only when a paused session resumes. Exploiting this, we place the hot set in HBM and the cold pool in HBF, forming a hot-cold KV hierarchy within the GPU memory tier. On agentic workloads with Qwen3-Coder-30B-A3B, our design delivers 14 ms time-between-tokens (TBT) and adds only $\approx$0.1 ms of resume latency on top of prefill, while hosting $24\times$ more concurrent sessions per GPU. By confining HBM to the hot set, our design also cuts read power by 7.6 kW per 8-GPU node relative to serving all KV from flash-establishing HBF as a cold-tier complement to HBM rather than its replacement.

补充信息

↑