arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

探究大语言模型推理在上下文长度与注意力架构上的能量缩放规律

Understanding the Energy Scaling of Large Language Model Inference Across Context Lengths and Attention Architectures

Molka Chkir, Syed Muhammad Danish, Jos Höll, Arghavan Asad

arXiv 2608.25096首次发表:更新:

发表机构

Algoma University; Reutlingen University(阿尔戈马大学; 罗伊特林根大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文通过实证研究,明确注意力架构是影响LLM解码能耗随上下文长度缩放的核心因素,发现GQA及带SWA的GQA能效更优,批处理可降低能耗与延迟,为LLM配置提供实用指导。

AI 中文摘要

大语言模型(LLMs)的日益普及引发了人们对其推理阶段能耗及环境影响的日益关注。本文对采用多头注意力(MHA)、分组查询注意力(GQA)及带滑动窗口注意力(SWA)的分组查询注意力的代表性开源LLMs的解码阶段能耗开展了系统性实证研究,以明确注意力架构如何在不同推理负载下影响解码阶段能耗。我们评估了四种模型在不同上下文长度、批大小及生成负载下的表现,同时利用NVIDIA硬件计数器测量GPU能耗,考察上下文长度、注意力机制、键值(KV)缓存增长及批处理对解码阶段能耗的影响。结果表明,注意力机制是决定解码能耗如何随上下文长度缩放的主要因素:MHA模型的能耗增长远比GQA模型陡峭,而带SWA的GQA模型则能维持近乎恒定的能耗。我们进一步发现,模型规模主要决定绝对能耗,而批处理可将每生成token的能耗及请求延迟降低最多87%。这些发现为选择高能效LLMs架构及推理配置提供了实用指导。

英文摘要

The growing adoption of large language models (LLMs) has raised increasing concerns about the energy consumption and environmental impact of inference. This paper presents a systematic empirical study of decode-phase energy consumption across representative open-source LLMs employing Multi-Head Attention (MHA), Grouped Query Attention (GQA), and Grouped Query Attention with Sliding Window Attention (SWA) to characterize how attention architecture influences decode-phase energy consumption under varying inference workloads. We evaluate four models across different context lengths, batch sizes, and generation workloads while measuring GPU energy using NVIDIA hardware counters. We examine the effects of context length, attention mechanism, Key-Value (KV) cache growth, and batching on decode-phase energy consumption. Results show that attention mechanism is the primary factor governing how decode energy scales with context length. MHA models exhibit substantially steeper energy growth than GQA models, whereas GQA with SWA maintains nearly constant energy consumption. We further show that model size primarily determines absolute energy consumption, while batching reduces both energy per generated token and request latency by up to 87%. These findings provide practical guidance for selecting energy-efficient LLM architectures and inference configurations.

Comments8 pages, 7 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑