发表机构
University of Southern California(南加州大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出混合三级框架HYMELL,结合分析建模与机器学习,在NVIDIA H100 GPU上对LLaMA 3 8B的预填充、解码阶段误差均低于5%,可实现LLM推理延迟与能耗的精准估计及能效优化。
AI 中文摘要
大语言模型(LLM)的快速扩展大幅提升了计算成本、能耗与推理延迟,精准估计对可持续人工智能部署及硬件感知设计至关重要。本研究提出LLM能耗与延迟混合建模框架HYMELL,这是一种混合三级框架,通过结合分析建模与机器学习(ML)来估计LLM的推理延迟与能耗。HYMELL通过三级层级对LLM执行过程建模:对基础操作进行分析估计,对高层组件进行ML预测,以及一个捕获预填充和解码阶段系统级开销的端到端模型。该框架支持多种架构,包括密集型与混合专家(MoE)前馈网络(FFN),以及多头注意力(MHA)与分组查询注意力(GQA)机制。在NVIDIA H100图形处理器(GPU)上评估时,HYMELL达到了高预测精度;值得注意的是,对于LLaMA 3 8B,它在预填充和解码阶段的误差均低于5%。通过直接根据架构参数预测执行成本,该框架支持快速、无硬件的设计空间探索与能效优化。
英文摘要
The rapid scaling of Large Language Models (LLMs) has significantly increased computational cost, energy consumption, and inference latency, making accurate estimation essential for sustainable artificial intelligence deployment and hardware-aware design. In this work, we introduce Hybrid Modeling for Energy and Latency of LLMs (HYMELL), a hybrid three-level framework for estimating LLM inference latency and energy by combining analytical modeling with machine learning (ML). HYMELL models LLM execution through a three-level hierarchy: analytical estimation of primitive operations, ML prediction of higher-level components, and an end-to-end model that captures system-level overheads across both prefill and decode phases. The framework supports diverse architectures, including dense and mixture-of-experts (MoE) feed-forward networks (FFNs), as well as multi-head attention (MHA) and grouped-query attention (GQA) mechanisms. Evaluated on an NVIDIA H100 graphics processing unit (GPU), HYMELL achieves high predictive accuracy; notably, for LLaMA 3 8B, it attains less than 5% error for both prefill and decode phases. By predicting execution costs directly from architectural parameters, it enables fast, hardware-free design space exploration and energy-efficient optimization.