arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FLINT:高效利用高带宽闪存实现容量可扩展的大语言模型推理加速

FLINT: Efficiently Leveraging High Bandwidth Flash for Capacity-Scalable LLM Inference Acceleration

Geraldo F. Oliveira, Arash Tavakkol, Xiangyu Zhu, Ahmet Caner Yüzügüler, Vamanan Arulchelvan, Lukas Cavigelli, Renzo Andri, Mohammad Sadrosadati, Jia Xinglei, Onur Mutlu, Zhou Ke, Shai Bergman, Ji Zhang

arXiv 2608.25062首次发表:更新:

发表机构

Huawei Technologies Switzerland AG; Huawei Technologies Co., Ltd.; ETH Zürich; HUST(华为技术瑞士有限公司; 华为技术有限公司; 苏黎世联邦理工学院; 华中科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

FLINT是一种工作负载驱动型HBF载体,通过三种机制解决现有HBF方案的三大挑战,实现高效利用HBF加速容量可扩展的LLM推理。

AI 中文摘要

大语言模型(LLM)推理日益受到加速器内存容量而非计算吞吐量的限制,这一约束在单加速器和小节点推理系统中尤为突出,有限的封装内内存容量限制了可部署模型的规模。HBF是一种新兴的3D堆叠NAND闪存技术,可提供数TB级的近加速器容量,是存储LLM权重的理想容量层。然而,现有的基于HBF的方案面临三大应用挑战:其一,依赖粗粒度的静态预取LLM权重,旨在隐藏NAND闪存设备微秒级的读取延迟,同时最大化HBF的读取吞吐量;其二,将NAND闪存管理任务(如刷新操作)暴露于加速器可见的关键推理路径;其三,未能针对工作负载行为对闪存管理机制进行专业化优化的机会。本文的目标是设计一种高效的HBF载体,将HBF作为内存容量层与HBM集成,同时解决上述三大挑战。为此,我们提出FLINT,一种用于容量可扩展LLM推理的工作负载驱动型HBF载体。FLINT引入三种机制:其一,硬件突发缓冲区控制器,动态合并并流水线化HBF读取,旨在利用现有NAND闪存缓冲区,同时维持高HBF带宽;其二,幻影平面刷新机制,通过低成本资源复制将与刷新相关的NAND闪存操作移至读取前台之外,从而将刷新操作从关键推理路径中移除;其三,只读FTL,用紧凑表替代SSD级的任意写入支持,该表用于将逻辑权重突发转换为物理HBF位置。

英文摘要

LLM inference is increasingly constrained by accelerator memory capacity rather than compute throughput. This constraint is especially acute in single-accelerator and small-node inference systems, where limited on-package memory capacity restricts the size of deployable models. HBF is an emerging 3D-stacked NAND flash technology that provides multi-terabyte near-accelerator capacity, making it a promising capacity tier for storing LLM weights. However, existing HBF-based proposals face three adoption challenges: they (1) rely on coarse-grained static prefetching for LLM weights aiming to hide the microsecond-level read latency of the NAND flash device while maximizing HBF's read throughput, (2) expose NAND flash management tasks (e.g., refresh operations) to the accelerator-visible critical inference path, and (3) miss optimization opportunities to specialize and optimize the flash-management mechanisms to the workload behavior. Our goal is to design an efficient HBF substrate that integrates HBF as a memory-capacity tier alongside HBM while addressing these three challenges. To this end, we propose FLINT, a workload-driven HBF substrate for capacity-scalable LLM inference. FLINT introduces three mechanisms: (1) a hardware burst-buffer controller that dynamically coalesces and pipelines HBF reads aiming to utilize existing NAND flash buffers while sustaining high HBF bandwidth, (2) a phantom-plane refresh mechanism, which removes refresh from the critical inference path by moving refresh-related NAND flash operations outside the read foreground back via low-cost resource duplication, and (3) a read-only FTL, which replaces SSD-class support for arbitrary writes with a compact table that translates logical weight bursts to physical HBF locations.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑