发表机构
Huawei Technologies Switzerland AG; Huawei Technologies Co., Ltd.; ETH Zürich; HUST(华为技术瑞士有限公司; 华为技术有限公司; 苏黎世联邦理工学院; 华中科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
FLINT是一种工作负载驱动型HBF载体,通过三种机制解决现有HBF方案的三大挑战,实现高效利用HBF加速容量可扩展的LLM推理。
AI 中文摘要
大语言模型(LLM)推理日益受到加速器内存容量而非计算吞吐量的限制,这一约束在单加速器和小节点推理系统中尤为突出,有限的封装内内存容量限制了可部署模型的规模。HBF是一种新兴的3D堆叠NAND闪存技术,可提供数TB级的近加速器容量,是存储LLM权重的理想容量层。然而,现有的基于HBF的方案面临三大应用挑战:其一,依赖粗粒度的静态预取LLM权重,旨在隐藏NAND闪存设备微秒级的读取延迟,同时最大化HBF的读取吞吐量;其二,将NAND闪存管理任务(如刷新操作)暴露于加速器可见的关键推理路径;其三,未能针对工作负载行为对闪存管理机制进行专业化优化的机会。本文的目标是设计一种高效的HBF载体,将HBF作为内存容量层与HBM集成,同时解决上述三大挑战。为此,我们提出FLINT,一种用于容量可扩展LLM推理的工作负载驱动型HBF载体。FLINT引入三种机制:其一,硬件突发缓冲区控制器,动态合并并流水线化HBF读取,旨在利用现有NAND闪存缓冲区,同时维持高HBF带宽;其二,幻影平面刷新机制,通过低成本资源复制将与刷新相关的NAND闪存操作移至读取前台之外,从而将刷新操作从关键推理路径中移除;其三,只读FTL,用紧凑表替代SSD级的任意写入支持,该表用于将逻辑权重突发转换为物理HBF位置。
英文摘要
LLM inference is increasingly constrained by accelerator memory capacity rather than compute throughput. This constraint is especially acute in single-accelerator and small-node inference systems, where limited on-package memory capacity restricts the size of deployable models. HBF is an emerging 3D-stacked NAND flash technology that provides multi-terabyte near-accelerator capacity, making it a promising capacity tier for storing LLM weights. However, existing HBF-based proposals face three adoption challenges: they (1) rely on coarse-grained static prefetching for LLM weights aiming to hide the microsecond-level read latency of the NAND flash device while maximizing HBF's read throughput, (2) expose NAND flash management tasks (e.g., refresh operations) to the accelerator-visible critical inference path, and (3) miss optimization opportunities to specialize and optimize the flash-management mechanisms to the workload behavior. Our goal is to design an efficient HBF substrate that integrates HBF as a memory-capacity tier alongside HBM while addressing these three challenges. To this end, we propose FLINT, a workload-driven HBF substrate for capacity-scalable LLM inference. FLINT introduces three mechanisms: (1) a hardware burst-buffer controller that dynamically coalesces and pipelines HBF reads aiming to utilize existing NAND flash buffers while sustaining high HBF bandwidth, (2) a phantom-plane refresh mechanism, which removes refresh from the critical inference path by moving refresh-related NAND flash operations outside the read foreground back via low-cost resource duplication, and (3) a read-only FTL, which replaces SSD-class support for arbitrary writes with a compact table that translates logical weight bursts to physical HBF locations.