arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

StreamDQ:用于可扩展人工智能推理加速的定制HBM中的近内存权重反量化

StreamDQ: Near-Memory Weight DeQuantization in Custom HBM for Scalable AI Inference Acceleration

Minki Jeong, Daegun Yoon, Soohong Ahn, Seungyong Lee, Nameun Kang, Hyeonseok Ju, Ieryung Park, Joonseop Sim, Youngpyo Joo, Hoshik Kim

arXiv 2607.08993首次发表:更新:

AI 中文总结

研究针对大语言模型内存和计算需求增长,提出StreamDQ架构,将反量化块集成到HBM基础芯片,通过内存端反量化减少GPU开销,实现高吞吐量推理,在混合精度GEMM和端到端LLM推理中均有显著性能提升。

AI 中文摘要

随着大语言模型(LLMs)规模的扩大,其内存和计算需求大幅增长,仅权重量化成为广泛采用的技术以减少模型大小并最小化精度损失。然而,在当前GPU上,基于CUDA核心的反量化会带来大量指令开销、片上流量和流水线停顿,成为高吞吐量、云规模LLM服务的主要瓶颈。为解决这些限制,我们提出StreamDQ,一种轻量级架构增强,可在内存子系统中进行实时反量化以实现高吞吐量、大批量LLM推理。StreamDQ将紧凑的反量化块(DQBs)集成到高带宽内存(HBM)的基础芯片中,并对标准内存加载执行内联反量化。每个内存读取请求上的轻量级边带标签选择反量化模式,同时保留传统加载语义。通过将反量化转移到内存端,StreamDQ消除了基于GPU侧CUDA核心的反量化,从而减少了GPU上的片上流量,并避免了大批量时反量化权重的额外HBM回写和重新加载。我们的评估表明,对于混合精度GEMM,StreamDQ实现了高达7.08倍的加速和90.23%的更低能耗,在12纳米CMOS工艺中每个DQB仅增加0.127平方毫米的面积和0.355瓦的功率开销。对于端到端LLM推理,StreamDQ将延迟降低了高达54.68%,并将解码吞吐量提高了高达2.20倍。

英文摘要

As large language models (LLMs) scale, their memory and computation demands have grown substantially, making weight-only quantization a widely adopted technique for reducing model size with minimal accuracy loss. However, on current GPUs, CUDA-core-based dequantization introduces substantial instruction overhead, on-chip traffic, and pipeline stalls, making it a major bottleneck for high-throughput, cloud-scale LLM serving. To address these limitations, we propose StreamDQ, a lightweight architectural enhancement that enables on-the-fly dequantization in the memory subsystem for high-throughput, large-batch LLM inference. StreamDQ integrates compact DeQuantization Blocks (DQBs) into the base die of high-bandwidth memory (HBM) and performs inline dequantization on standard memory loads. A lightweight sideband tag on each memory read request selects the dequantization mode while preserving conventional load semantics. By relocating dequantization to the memory side, StreamDQ eliminates GPU-side CUDA-core-based dequantization, thereby reducing on-chip traffic on the GPU and avoiding extra HBM write-back and reload of dequantized weights at large batch sizes. Our evaluation shows that StreamDQ achieves up to 7.08$\times$ speedup and 90.23\% lower energy for mixed-precision GEMM, with only 0.127\,mm$^2$ area and 0.355\,W power overhead per DQB in a 12\,nm CMOS process. For end-to-end LLM inference, StreamDQ reduces latency by up to 54.68\% and improves decode throughput by up to 2.20$\times$.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑