AI 中文总结
研究针对大语言模型内存和计算需求增长,提出StreamDQ架构,将反量化块集成到HBM基础芯片,通过内存端反量化减少GPU开销,实现高吞吐量推理,在混合精度GEMM和端到端LLM推理中均有显著性能提升。
AI 中文摘要
随着大语言模型(LLMs)规模的扩大,其内存和计算需求大幅增长,仅权重量化成为广泛采用的技术以减少模型大小并最小化精度损失。然而,在当前GPU上,基于CUDA核心的反量化会带来大量指令开销、片上流量和流水线停顿,成为高吞吐量、云规模LLM服务的主要瓶颈。为解决这些限制,我们提出StreamDQ,一种轻量级架构增强,可在内存子系统中进行实时反量化以实现高吞吐量、大批量LLM推理。StreamDQ将紧凑的反量化块(DQBs)集成到高带宽内存(HBM)的基础芯片中,并对标准内存加载执行内联反量化。每个内存读取请求上的轻量级边带标签选择反量化模式,同时保留传统加载语义。通过将反量化转移到内存端,StreamDQ消除了基于GPU侧CUDA核心的反量化,从而减少了GPU上的片上流量,并避免了大批量时反量化权重的额外HBM回写和重新加载。我们的评估表明,对于混合精度GEMM,StreamDQ实现了高达7.08倍的加速和90.23%的更低能耗,在12纳米CMOS工艺中每个DQB仅增加0.127平方毫米的面积和0.355瓦的功率开销。对于端到端LLM推理,StreamDQ将延迟降低了高达54.68%,并将解码吞吐量提高了高达2.20倍。
英文摘要
As large language models (LLMs) scale, their memory and computation demands have grown substantially, making weight-only quantization a widely adopted technique for reducing model size with minimal accuracy loss. However, on current GPUs, CUDA-core-based dequantization introduces substantial instruction overhead, on-chip traffic, and pipeline stalls, making it a major bottleneck for high-throughput, cloud-scale LLM serving. To address these limitations, we propose StreamDQ, a lightweight architectural enhancement that enables on-the-fly dequantization in the memory subsystem for high-throughput, large-batch LLM inference. StreamDQ integrates compact DeQuantization Blocks (DQBs) into the base die of high-bandwidth memory (HBM) and performs inline dequantization on standard memory loads. A lightweight sideband tag on each memory read request selects the dequantization mode while preserving conventional load semantics. By relocating dequantization to the memory side, StreamDQ eliminates GPU-side CUDA-core-based dequantization, thereby reducing on-chip traffic on the GPU and avoiding extra HBM write-back and reload of dequantized weights at large batch sizes. Our evaluation shows that StreamDQ achieves up to 7.08$\times$ speedup and 90.23\% lower energy for mixed-precision GEMM, with only 0.127\,mm$^2$ area and 0.355\,W power overhead per DQB in a 12\,nm CMOS process. For end-to-end LLM inference, StreamDQ reduces latency by up to 54.68\% and improves decode throughput by up to 2.20$\times$.