发表机构
Georgia State University; DEVCOM Army Research Laboratory(佐治亚州立大学; 陆军研究实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对异构边缘设备上大语言模型的延迟预测问题,提出运行时感知延迟预测框架,将推理请求配置化,分阶段处理并融合数据,经实验验证该框架能降低分析成本,支持跨平台部署。
AI 中文摘要
准确的延迟预测对于在异构边缘设备上部署大语言模型至关重要,因为推理延迟受多种因素影响。本文提出了一个面向部署的大语言模型选择的运行时感知延迟预测框架。该框架将每个推理请求表示为硬件 - 运行时 - 模型 - 提示配置,将推理分为预填充和解码阶段,并通过门控预测模型自适应地融合静态描述符和动态硬件遥测数据。我们使用Pixel移动设备评估了该框架,并在Jetson Nano、Orange Pi 5 Pro和RTX 3090级GPU平台上验证了分析管道。结果表明,通过轻量级校准的运行时感知预测可以降低分析成本,并支持跨异构边缘平台的延迟感知大语言模型部署。
英文摘要
Accurate latency prediction is critical for deploying large language models (LLMs) on heterogeneous edge devices, where inference latency is affected by model architecture, prompt behavior, runtime backend, hardware utilization, dynamic voltage and frequency scaling (DVFS), and thermal variation. This paper presents a runtime-aware latency prediction framework for deployment-oriented LLM selection. The framework represents each inference request as a hardware-runtime-model-prompt configuration, separates inference into prefill and decode phases, and adaptively fuses static descriptors with dynamic hardware telemetry through a gated prediction model. We evaluate the framework using Pixel mobile devices and validate the profiling pipeline on Jetson Nano, Orange Pi 5 Pro, and an RTX 3090-class GPU platform. On Pixel 8, the full predictor improves total-latency R-squared from 0.953 to 0.960 and decode-latency R-squared from 0.957 to 0.973 over a static-only baseline. On Pixel 8 Pro, it improves prefill-latency R-squared from -1.383 to 0.966. For cross-device transfer, calibration improves Pixel 8 Pro to Pixel 8 total-latency R-squared from -0.974 to 0.940 and decode-latency R-squared from -1.085 to 0.927. Heterogeneous profiling further shows that latency is highly device- and runtime-dependent: the same SmolLM2 model family reaches 8.42 tokens/s on Orange Pi 5 Pro but 64.38 tokens/s on an RTX 3090-class GPU. These results demonstrate that runtime-aware prediction with lightweight calibration can reduce profiling cost and support latency-aware LLM deployment across heterogeneous edge platforms.