arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.29575cs.DCcs.PF

SLIM:面向大语言模型服务的饱和度感知轻量级性能建模

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving

Pol G. Recasens, Ferran Agullo, Yue Zhu, Chen Wang, Jordi Torres, Josep Ll. Berral

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对LLM服务吞吐量饱和问题,通过GPU剖析明确其源于解码阶段注意力内核的DRAM带宽饱和,提出SLIM半分析性能模型及BCA工具,可高效预测吞吐量与延迟并优化批处理配置。

中文摘要 AI 辅助

大语言模型(LLM)服务通常通过增大批量大小来提升吞吐量,但性能最终会达到与部署环境相关的平台期,超过该平台期后,更大的批量仅能带来微弱的增益,却会增加延迟和GPU内存消耗。此前的研究将该行为归因于高带宽内存(HBM)/动态随机存取存储器(DRAM)带宽限制,但根本原因主要由概念性论证或高层次性能观察支撑。作为首个贡献,我们采用硬件剖析技术开展了详细的GPU表征,证明吞吐量饱和源于解码阶段的注意力内核。具体而言,我们表明,随着活跃上下文长度增加(而非仅批量大小增大),注意力内核近乎恒定的算术强度会驱动DRAM带宽饱和,而实现的计算吞吐量仍远低于硬件极限。基于该分析,我们提出了批处理配置顾问(BCA),其可选择满足目标延迟约束的最高吞吐量批处理配置,并针对评估的OPT模型,识别出最多55GB的GPU内存分配可在吞吐量损失极小的情况下避免。为支持这些建议,我们引入了SLIM(Saturation-Aware Lightweight Performance Model,饱和度感知轻量级性能模型),这是一种半分析模型,可通过Transformer计算和内存流量的分析公式预测LLM推理的吞吐量与延迟。在评估场景中,SLIM的表现优于代表性性能建模基线,且能成功泛化到此前未见过的运行条件。

英文摘要

Large language model (LLM) serving commonly increases batch size to improve throughput, but performance eventually reaches a deployment-dependent plateau beyond which larger batches provide marginal gains while increasing latency and GPU memory consumption. Previous studies have attributed this behavior to HBM/DRAM bandwidth limitations, but the underlying causes have primarily been supported by conceptual arguments or high-level performance observations. As our first contribution, we present a detailed GPU characterization using hardware profiling techniques, demonstrating that throughput saturation originates in the attention kernels during the decode phase. Specifically, we show that their nearly constant arithmetic intensity as active-context lengths increases -not merely larger batch sizes- drives DRAM-bandwidth saturation, while the achieved compute throughput remains far below the hardware limit. Building on this analysis, we present the Batching Configuration Advisor (BCA), which selects the highest-throughput batching configuration satisfying a target latency constraint and identifies up to 55 GB of GPU memory allocation that can be avoided for the evaluated OPT models with minimal throughput loss. To enable these recommendations, we introduce SLIM (Saturation-Aware Lightweight Performance Model), a semi-analytical model that predicts LLM inference throughput and latency from analytical formulations of Transformer computation and memory traffic. Across the evaluated scenarios, SLIM outperforms representative performance-modeling baselines while successfully generalizing to previously unseen operating conditions.

补充信息

↑