AI 中文总结
该研究针对异构DRAM-PIM-GPU系统,通过系统评估OPT、Mamba2系列模型,揭示了三项设计原则,为内存加速LLM系统架构设计提供了指导。
AI 中文摘要
基于DRAM的异构内存处理(PIM)-GPU系统有望为解码阶段的大语言模型(LLM)推理带来显著的效率提升,尤其在长输出生成场景中,但当前设计实践忽略了决定实际性能的关键因素。通过对多种架构和工作负载(OPT-7B/70B、Mamba2-2.7B/70B)的系统评估,研究揭示了三项基本设计原则:(i)静态功耗(DRAM泄漏、刷新操作及GPU空闲功耗)可主导效率计算,在实际部署中,仅考虑动态功耗的模型会高估token/s/W指标,对于Mamba2-2.7B、批次大小1、128个输入token及2048个输出token的场景,高估倍数可达3.85倍;(ii)在所有评估模型和工作负载中,解码性能随通道数单调非递减,低批次工作负载通常在高通道数下达到性能平台;在固定容量扫描下,所有模型共享一个近似最优的层级配置,基于注意力的模型若配置不当会承受大得多的性能损失;(iii)工作负载映射策略带来的提升有限,内核级延迟和功耗的降低分别最高为14.0%和17.4%,端到端提升最高为5.6%,并非主要瓶颈。显著的效率提升需要全系统协同优化,这些原则为设计下一代内存加速LLM系统的架构师提供了设计空间指导。
英文摘要
Heterogeneous DRAM-based processing-in-memory (PIM)-GPU systems promise significant efficiency gains for decode-phase large language model (LLM) inference, particularly in long-output generation, yet current design practices overlook critical factors that determine real-world performance. Through systematic evaluation of diverse architectures and workloads (OPT-7B/70B, Mamba2-2.7B/70B), we reveal three fundamental design principles: (i) static power consumption (DRAM leakage, refresh, and GPU idle power) can dominate the efficiency calculus, causing dynamic-only models to overestimate tokens/s/W by up to 3.85X for realistic deployments (Mamba2-2.7B, batch size 1, 128 input tokens, and 2,048 output tokens); (ii) decoding performance is monotonically non-decreasing with channel count across all evaluated models and workloads, generally plateauing at high channel counts for low-batch workloads; under a fixed-capacity sweep, all models instead share a common near-optimal hierarchy configuration, with substantially larger misconfiguration penalties for attention-based models; (iii) workload mapping strategies provide bounded improvements (up to 14.0%/17.4% kernel-level latency/energy reduction, up to 5.6% end-to-end gain) and are not primary bottlenecks. Significant efficiency gains require system-wide co-optimization. These principles provide design-space guidance for architects designing the next generation of memory-accelerated LLM systems.
CommentsAccepted at 2026 IEEE International System-on-Chip Conference (SOCC)