arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.13127cs.AR

HBF在大语言模型(LLM)服务系统中的潜在应用

Potential Applications of HBF in LLM Serving Systems

Yihan Yin, Yinlun Zhao, Zhixin Yun, Guanying Wu, Feng Zhu, Kai Tao, Shu Li, Fei Huang, Zhe Zhang, Shuangchen Li, Hongzhong Zheng

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对LLM服务的内存容量限制,提出将HBF集成到GPU内存层级的方案,通过仿真验证其可提升MoE和多模型服务性能,需保留HBM驻留执行路径。

中文摘要 AI 辅助

随着模型权重、键值缓存(KV caches)以及所服务模型变体数量的不断增长,LLM服务日益受到内存容量的限制。本报告研究了高带宽闪存(High-Bandwidth Flash,HBF)作为基于高带宽内存(HBM)的服务系统的容量导向扩展方案。我们首先探讨如何将HBF集成到GPU内存层级中,同时不损害计算核心所需的带宽。接着,我们将新增容量的系统级价值建模为只读为主的模型状态对象的扩展驻留空间。基于此视角,HBF可通过支持更多专家副本提升混合专家模型(MoE)服务的性能,还可通过减少模型加载并支持热模型复制提升多模型服务的性能。我们的仿真结果表明,这些优势取决于在使用HBF扩展模型权重驻留集的同时,保留HBM驻留的执行路径。

英文摘要

LLM serving is increasingly constrained by memory capacity as model weights, KV caches, and the number of served model variants continue to grow. This report examines High-Bandwidth Flash (HBF) as a capacity-oriented extension to HBM-based serving systems. We first discuss how HBF can be integrated into the GPU memory hierarchy without undermining the bandwidth expected by the compute die. We then model the system-level value of added capacity as expanded residency for read-mostly model-state objects. Under this view, HBF can improve MoE serving by enabling more expert replicas and can improve multi-model serving by reducing model loading and supporting hot-model replication. Our simulation results show that these benefits depend on preserving the HBM-resident execution path while using HBF to expand the resident set of model weights.

↑