arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.30509cs.AR

CHIPSMORE:用于多模式多请求大语言模型推理加速的互连内及内存内小芯片

CHIPSMORE: Compute-in-Interconnect and -Memory Chiplets for Multi-Mode Multi-Request LLM Inference Acceleration

Yue Jiet Chong, Yimin Wang, Zhen Wu, Zixuan Wang, Wei Zhang, Xuanyao Fong

首次发表
浏览论文内容

中文总结 AI 辅助

CHIPSMORE是一款集成互连内计算与CIM的LLM推理加速器,通过异构处理单元、分层KV内存、非复制多请求管道等设计,在Mistral-7B上较Nvidia H100实现2.38倍吞吐量与27倍能效提升,消除多请求权重复制。

中文摘要 AI 辅助

大语言模型(LLM)推理在适配模式、上下文长度和请求并发度方面存在显著差异,这给内存内计算(CIM)加速器维持高利用率、内存效率和可扩展性能带来了挑战。本文提出了CHIPSMORE,一种多模式多请求LLM推理加速器,它集成了互连内计算和CIM,以支持不同工作负载下的基础模式和低秩适配(LoRA)推理。CHIPSMORE采用由阻变RAM模拟内存内计算(RRAM-ACIM)和静态RAM数字内存内计算(SRAM-DCIM)组成的异构处理单元,这些单元通过可编程的处理单元间计算网络(IPCN)互连。一种可组合的分层键值(KV)内存方案根据工作负载需求动态分配路由器暂存区、SRAM-DCIM和嵌入式DRAM(eDRAM)资源,从而实现对长上下文和批量推理的可扩展支持。此外,一种非复制的多请求执行管道利用请求级并行性,无需复制预训练权重;而状态感知资源重配置机制可选择性保留运行时状态并对非活跃资源进行电源门控,以提高能效。采用周期精确的软硬件协同仿真进行评估,结果表明CHIPSMORE在不同模型规模、上下文长度和批量大小下均能有效维持高吞吐量,同时保持良好的功率扩展能力。与Nvidia H100相比,CHIPSMORE在Mistral-7B推理上实现了高达2.38倍的吞吐量和27倍的能效提升,同时消除了多请求服务中的权重复制。

英文摘要

Large language model (LLM) inference exhibits substantial variability across adaptation modes, context lengths, and request concurrency, creating challenges for maintaining high utilization, memory efficiency, and scalable performance on compute-in-memory (CIM) accelerators. This paper presents CHIPSMORE, a multi-mode and multi-request LLM inference accelerator that integrates compute-in-interconnect and CIM to support both base-mode and low-rank adaptation (LoRA) inference under diverse workloads. CHIPSMORE employs heterogeneous processing elements consisting of resistive RAM analog compute-in-memory (RRAM-ACIM) and static RAM digital compute-in-memory (SRAM-DCIM) interconnected through a programmable Inter-PE computational network (IPCN). A composable hierarchical key-value (KV) memory scheme dynamically allocates router scratchpad, SRAM-DCIM, and embedded DRAM (eDRAM) resources according to workload requirements, enabling scalable support for long-context and batched inference. Furthermore, a non-replicated multi-request execution pipeline exploits request-level parallelism without duplicating pretrained weights, while a state-aware resource reconfiguration mechanism selectively retains runtime states and power-gates inactive resources to improve energy efficiency. Evaluation using cycle-accurate hardware-software co-simulation demonstrates that CHIPSMORE effectively sustains high throughput across varying model sizes, context lengths, and batch sizes while maintaining favorable power scaling. Compared with Nvidia H100, CHIPSMORE achieves up to $2.38\times$ higher throughput and $27\times$ higher energy efficiency on Mistral-7B inference while eliminating weight replication for multi-request serving.

发表机构

  • National University of Singapore(新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

↑