arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于混合分析-机器学习预测器的大语言模型推理延迟与能耗的多级建模

Multi-Level Modeling of Large Language Model Inference Latency and Energy via Hybrid Analytical--Machine-Learning Predictors

Saeid Shokoufa, Mohammad Erfan Sadeghi, Mehdi Kamal, Massoud Pedram

arXiv 2608.06723首次发表:更新:

发表机构

University of Southern California(南加州大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出混合三级框架HYMELL,结合分析建模与机器学习,在NVIDIA H100 GPU上对LLaMA 3 8B的预填充、解码阶段误差均低于5%,可实现LLM推理延迟与能耗的精准估计及能效优化。

AI 中文摘要

大语言模型(LLM)的快速扩展大幅提升了计算成本、能耗与推理延迟,精准估计对可持续人工智能部署及硬件感知设计至关重要。本研究提出LLM能耗与延迟混合建模框架HYMELL,这是一种混合三级框架,通过结合分析建模与机器学习(ML)来估计LLM的推理延迟与能耗。HYMELL通过三级层级对LLM执行过程建模:对基础操作进行分析估计,对高层组件进行ML预测,以及一个捕获预填充和解码阶段系统级开销的端到端模型。该框架支持多种架构,包括密集型与混合专家(MoE)前馈网络(FFN),以及多头注意力(MHA)与分组查询注意力(GQA)机制。在NVIDIA H100图形处理器(GPU)上评估时,HYMELL达到了高预测精度;值得注意的是,对于LLaMA 3 8B,它在预填充和解码阶段的误差均低于5%。通过直接根据架构参数预测执行成本,该框架支持快速、无硬件的设计空间探索与能效优化。

英文摘要

The rapid scaling of Large Language Models (LLMs) has significantly increased computational cost, energy consumption, and inference latency, making accurate estimation essential for sustainable artificial intelligence deployment and hardware-aware design. In this work, we introduce Hybrid Modeling for Energy and Latency of LLMs (HYMELL), a hybrid three-level framework for estimating LLM inference latency and energy by combining analytical modeling with machine learning (ML). HYMELL models LLM execution through a three-level hierarchy: analytical estimation of primitive operations, ML prediction of higher-level components, and an end-to-end model that captures system-level overheads across both prefill and decode phases. The framework supports diverse architectures, including dense and mixture-of-experts (MoE) feed-forward networks (FFNs), as well as multi-head attention (MHA) and grouped-query attention (GQA) mechanisms. Evaluated on an NVIDIA H100 graphics processing unit (GPU), HYMELL achieves high predictive accuracy; notably, for LLaMA 3 8B, it attains less than 5% error for both prefill and decode phases. By predicting execution costs directly from architectural parameters, it enables fast, hardware-free design space exploration and energy-efficient optimization.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑