arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.18675cs.AR

HBFlex:一种桥接细粒度LLM状态与粗粒度HBF并行执行的灵活内存系统

HBFlex: A Flexible Memory System for Bridging Fine-Grained LLM States and Coarse-Grained HBF Parallel Execution

  • Peking University(北京大学)
  • HKUST(香港科技大学)
  • Alibaba Group(阿里巴巴集团)

机构由 AI 辅助整理,请以论文原文为准。

Shuzhang Zhong, Weikai Xu, Yifan Zhou, Tongbin Zhao, Tenghao Zhao, Yifei Kang, Cunyin Chang, Shu Li, Guangyu Sun, Meng Li

AI总结:

针对LLM全HBF内存服务中KV读写与回收的挑战,提出HBFlex系统,通过平衡放置、聚合写回及生命周期引导回收,实现吞吐量较FlashAccel提升1.58倍、较H3提升3.30倍。

AI中文摘要:

大型语言模型(LLM)需要不断增加的内存容量以适应不断增长的模型权重和KV缓存。高带宽闪存(HBF)通过大规模平面级并行性提供高内存密度和聚合读带宽,使其成为LLM服务的一个有吸引力的选择。然而,完全从HBF服务LLM带来了三个挑战:细粒度的KV读取造成放置和访问不平衡,增量写入干扰前台读取,以及混合KV生命周期放大了垃圾回收。混合HBM/HBF设计保留HBM以支持动态KV管理,但这种分配在固定封装预算下减少了可用的HBF资源,限制了聚合HBF带宽。我们提出了HBFlex,一个全HBF内存系统,对KV读取、写入和回收进行协调优化。HBFlex平衡KV放置和注意力访问以提高平面利用率。它聚合增量更新并在足够长的计算窗口内调度回写以减少读写干扰。它还结合了生命周期引导的块打包与延迟回收以减少有效页迁移。我们通过不同配置下的轨迹驱动模拟评估HBFlex。HBFlex相对于FlashAccel实现了平均吞吐量加速比高达1.58倍,相对于H3实现了3.30倍,得益于更高的HBF带宽和更高效的动态KV缓存读取、写入和擦除管理。

英文摘要:

Large language models (LLMs) require increasing memory capacity to accommodate growing model weights and KV caches. High-Bandwidth Flash (HBF) offers high memory density and aggregate read bandwidth through massive plane-level parallelism, making it an attractive option for LLM serving. However, serving LLMs entirely from HBF introduces three challenges: fine-grained KV reads create placement and access imbalance, incremental writes interfere with foreground reads, and mixed KV lifetimes amplify garbage collection. Hybrid HBM/HBF designs retain HBM to support dynamic KV management, but this allocation reduces the HBF resources available under a fixed packaging budget, limiting aggregate HBF bandwidth. We present HBFlex, a full-HBF memory system with coordinated optimizations for KV reads, writes, and reclamation. HBFlex balances KV placement and attention accesses to improve plane utilization. It aggregates incremental updates and schedules writeback within sufficiently long compute windows to reduce write--read interference. It also combines lifetime-guided block packing with deferred reclamation to reduce valid-page migration. We evaluate HBFlex through trace-driven simulation across different configurations. HBFlex achieves average throughput speedups of up to 1.58$\times$ over FlashAccel and 3.30$\times$ over H3, benefiting from higher HBF bandwidth and more efficient management of dynamic KV-cache reads, writes, and erases.

↑