arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LLMET:支持新兴M3D存储器的跨层评估以实现高能效LLM服务

LLMET: Enabling Cross-Layer Evaluation of Emerging M3D Memories for Energy-Efficient LLM Serving

Ming-Yen Lee, Hanchen Yang, Faaiq Waqar, Harsono Simka, Tushar Krishna, Muhammed Ahosan Ul Karim, Shimeng Yu

arXiv 2607.26491首次发表:更新:

发表机构

Georgia Institute of Technology; Samsung Semiconductor Inc.(佐治亚理工学院; 三星半导体公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究开发跨层仿真框架LLMET,通过实验证实采用M3D技术扩展片上缓存可显著降低不同平台及场景下LLM服务的预填充、解码阶段能耗,为高能效LLM服务系统提供新方案。

AI 中文摘要

随着部署规模扩大,大型语言模型(LLM)服务的能耗正成为一项重大系统挑战,这由硬件功耗、热约束以及不断上涨的电力成本所驱动。芯片能耗的一个关键来源是有限的片上缓存与片外高带宽存储器(HBM)之间的数据传输。与此同时,诸如在逻辑芯片的后端(BEOL)集成缓存存储器的单片3D(M3D)集成等新兴存储器技术,能够实现更大、更密集的片上存储器,为减少成本高昂的片外数据传输创造了新机遇。然而,目前尚不清楚利用新兴技术持续扩展片上存储器是否能有效提升LLM服务的能效。为解决这一问题,我们开发了经验证的跨层仿真框架LLMET(LLM with Emerging Technology),并针对广泛的模型、应用和平台,开展了大容量片上存储器技术影响的综合研究。基于在双NVIDIA A100 GPU配置下的LLMET仿真,利用M3D技术将L2缓存从40MB扩展至1GB,可使Llama3.1-70B在16K上下文窗口的预填充阶段芯片能耗降低44%;在8个类NVIDIA B200平台上,将L2缓存从128MB扩展至4GB可使预填充能耗最高节省24%;对于边缘平台及工作负载,将8MB缓存扩展至256MB时,解码能耗可节省30%。这些结果凸显了超大规模片上存储器在构建高能效LLM服务系统方面的潜力。

英文摘要

The energy consumption of Large Language Model (LLM) serving is becoming a major system challenge as deployment scales, driven by hardware power and thermal constraints and rising electricity costs. A key contributor to chip energy dissipation is data movement between limited on-chip cache and off-chip High Bandwidth Memory (HBM). Meanwhile, emerging memory technologies such as monolithic 3D (M3D) integration of cache memories at the Back-End-Of-Line (BEOL) of logic chips enable larger and denser on-chip memories, creating new opportunities to reduce costly off-chip traffic. However, it remains unclear whether continuously scaling on-chip memory using emerging technologies can effectively improve the energy efficiency of LLM serving. To address this gap, we develop LLMET (LLM with Emerging Technology), a validated cross-layer simulation framework, and conduct a comprehensive study on the impact of large-capacity on-chip memory technologies across a broad range of models, applications and platforms. Utilizing M3D technology to expand the L2 cache from 40MB to 1GB yields a 44% reduction in chip energy during the Llama3.1-70B prefill phase with a 16K context window, based on LLMET simulation on a dual NVIDIA A100 GPU setup. On the 8x NVIDIA B200-like platform, extending the L2 cache from 128MB to 4GB saves the prefill energy by up to 24%. For the edge platform and workloads, the decode energy saving reaches 30% when increasing the 8MB cache size to 256MB. These results highlight the promise of ultra-large on-chip memories for energy-efficient LLM serving systems.

Comments6 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑