arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.00857cs.AR

基于平衡数据流与细粒度并行IMC-NoC架构的大语言模型推理

LLM Inference on IMC-NoC Architecture with Balanced Dataflow and Fine-Grained Parallelism

  • National University of Singapore(新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

Yimin Wang, Yue Jiet Chong, Xuanyao Fong

AI总结:

该研究提出软硬件协同设计的LEAP架构,通过预填充-解码拆分方法优化LLM推理,较商用GPU实现了吞吐量≥1.52倍、能效24.91倍的提升。

AI中文摘要:

大语言模型(LLM)推理已成为一项关键服务,但其对内存带宽、计算密度和通信效率提出了前所未有的要求。内存计算(IMC)是缓解内存墙问题的有前景方案,但LLM的异构数据动态性需要补充资源来处理运行时生成的中间数据。此外,LLM的海量参数需要可扩展架构,而片上数据移动往往是主要性能瓶颈。本文提出一种软硬件协同设计框架,将分布式计算、内存和通信统一为无缝的处理-通信结构。硬件方面,我们提出名为LEAP的可扩展架构,集成IMC处理单元(PE)、NMC PE和INC,使每个硬件层执行专门任务:IMC处理静态权重,NMC处理动态数据,INC处理部分结果归约。软件方面,我们引入针对LLM服务关键指标(包括吞吐量和延迟)优化的分区、映射与调度框架。为解决预填充和解码阶段不同的计算强度,我们提出预填充-解码拆分方法,动态重配置PE组织以最大化资源利用率。与商用GPU平台相比,所提架构的吞吐量提升≥1.52倍,能效提升24.91倍。

英文摘要:

LLM inference has become an essential service, yet it imposes unprecedented demands on memory bandwidth, computational density, and communication efficiency. While IMC is a promising solution to the memory wall issue, the heterogeneous data dynamicity of LLM requires complementary resources to handle intermediate data generated during run-time. Furthermore, the massive number of parameters in LLM necessitates scale-up architectures where on-chip data movement is often the primary performance bottleneck. This article presents a hardware-software co-design framework that unifies distributed compute, memory, and communication into a seamless processing-communication fabric. On the hardware side, we propose a scalable architecture, named LEAP, that integrates IMC PE, NMC PE, and INC. This allows each hardware layer to execute specialized tasks: IMC for static weights, NMC for dynamic data, and INC for partial result reduction. On the software side, we introduce a partitioning, mapping, and scheduling framework optimized for key metrics in LLM serving, including throughput and latency. To address the distinct computational intensities of the prefill and decode phases, we present a prefill-decode disaggregation approach that dynamically reconfigures PE organizations to maximize resource utilization. Compared to commercial GPU platforms, the proposed architecture provides a throughput and an energy efficiency improvement of $\geq{}1.52\times$ and $24.91\times$, respectively.

补充信息

↑