arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.01891cs.DCcs.AI

通过解耦注意力机制与前馈网络及灵活频率缩放实现高能效大语言模型服务

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling

Cunchen Hu, Liangliang Xu, Tian Liu, Min Lyu, Yongkun Li, Sa Wang, Shuo Quan, Yanan Yang, Wenda Tang, Yiduo Wang, Fu Yu, Jie Wu

首次发表
浏览论文内容

中文总结 AI 辅助

AFlex框架通过全局调度器、局部DVFS控制器及交错A/F流水线,优化解耦A/F服务的资源配置与GPU频率,在满足SLO的同时将每token能耗降低最高49%

中文摘要 AI 辅助

大语言模型(LLM)服务覆盖各类对服务水平目标(SLO)有严格要求的应用,常需GPU以最高频率运行,导致能耗上升。现有能耗管理方法仅在请求或推理阶段层面调整GPU频率,忽略了注意力机制(Attention)与前馈网络(FFN)在算子层面的频率敏感性差异。研究发现,注意力机制与前馈网络(A/F)的能耗最优频率存在差异,且随推理阶段、工作负载和系统配置变化。然而,运行时的可变性及A/F独立频率控制会产生巨大搜索空间和高通信开销。为应对这些挑战,本文提出AFlex框架,用于联合优化解耦A/F服务的资源配置与GPU频率缩放。AFlex引入全局调度器和局部算子级动态电压频率缩放(DVFS)控制器,以确定A/F的资源分配和频率;还引入带动态微批量深度与自适应请求批处理的交错A/F流水线,减少流水线气泡。本文在SGLang中实现AFlex,并基于NVIDIA A800 GPU,采用Qwen3-32B和Mixtral-8×7B模型,在生产级对话与编码轨迹上进行评估。结果显示,AFlex在满足首 token 时间(TTFT)和每 token 输出时间(TPOT)SLO的前提下,相比最先进的解耦服务,每 token 能耗降低达49%;相比频率缩放系统,每 token 能耗降低达48%。

英文摘要

Large language model (LLM) serving spans diverse applications with stringent service-level objectives (SLOs), often requiring GPUs to run at maximum frequencies and increasing energy consumption. Existing energy-management approaches adapt GPU frequencies only at the request or inference-phase level, overlooking operator-level differences in frequency sensitivity between Attention and feed-forward networks (FFNs). We find that the energy-optimal frequencies of Attention and FFN (A/F) differ and vary with the inference phase, workload, and system configurations. However, runtime variability and independent A/F frequency control create a large search space and high communication overhead. To address these challenges, we present AFlex, a framework that jointly optimizes resource provisioning and GPU frequency scaling for disaggregated A/F serving. AFlex introduces a global scheduler and a local operator-level dynamic voltage and frequency scaling (DVFS) controller to determine A/F resource allocations and frequencies. It further introduces an interleaved A/F pipeline with dynamic microbatch depth and adaptive request batching to reduce pipeline bubbles. We implement AFlex in SGLang and evaluate it on NVIDIA A800 GPUs using Qwen3-32B and Mixtral-8$\times$7B under production Conversation and Coding traces. \AFlex reduces energy per token by up to 49\% over state-of-the-art disaggregated serving and 48\% over frequency-scaling systems while satisfying TTFT and TPOT SLOs.

发表机构

  • China Telecom Cloud Computing Research Institute(中国电信云计算研究院)
  • Xidian University(西安电子科技大学)
  • University of Science and Technology of China(中国科学技术大学)

机构由 AI 辅助整理,请以论文原文为准。

↑