LLM推理在边缘连续体硬件上的权衡测量研究
A Measurement Study of LLM Inference Trade-offs Across Edge Continuum Hardware
浏览论文内容
中文总结 AI 辅助
本研究测量了边缘连续体上LLM推理的权衡,发现GPU服务器延迟最低、Jetson Orin能耗更低,且参数大小不能预测性能,需考虑流式开销优化部署。
中文摘要 AI 辅助
大型语言模型(LLM)越来越多地被用作智能Web服务的后端,但在边缘连续体上提供服务需要在质量、延迟、模型占用空间和能量之间取得平衡。本文对自托管LLM推理在边缘和近边缘部署节点上进行了受控测量研究:一个NVIDIA Jetson AGX Orin和一个具有仅CPU和GPU启用推理模式的近边缘服务器。我们使用固定的问答工作负载评估了多个开放权重LLM和量化变体,并将它们与作为云托管准确性和延迟参考的GPT-4o进行比较。我们的基准测试流程报告了准确性、模型占用空间、每令牌解码延迟、预填充延迟和总体执行能量。结果表明,启用GPU的服务器执行提供了最低的计算侧延迟,而Jetson Orin显示出较低的测量能量,这与我们设置下其较低的平台功耗一致。仅CPU执行在我们的工作负载中始终在延迟方面处于劣势,并显示出较高的测量能量。我们还表明,参数数量和下载的权重文件大小单独并不能可靠地预测观察到的准确性或延迟。最后,使用帕累托前沿分析,我们研究了在可能的流式令牌传递开销下部署决策可能如何变化,强调仅计算侧推理指标可能导致对延迟敏感的交互式Web服务的次优放置。
英文摘要
Large language models (LLMs) are increasingly used as backends for intelligent web services, but serving them across the edge continuum requires balancing quality, latency, model footprint, and energy. This paper presents a controlled measurement study of self-hosted LLM inference across edge and near-edge deployment nodes: an NVIDIA Jetson AGX Orin and a near-edge server with CPU-only and GPU-enabled inference modes. We evaluate multiple open-weight LLMs and quantization variants using a fixed question-answering workload, and compare them against GPT-4o as a cloud-hosted accuracy and latency reference. Our benchmarking pipeline reports accuracy, model footprint, per-token decoding latency, prefill latency, and overall execution energy. The results show that GPU-enabled server execution provides the lowest compute-side latency, while Jetson Orin shows lower measured energy, consistent with its lower platform power under our setup. CPU-only execution is consistently dominated in latency for our workload and shows higher measured energy. We also show that parameter count and downloaded weight-file size alone do not reliably predict observed accuracy or latency. Finally, using Pareto-frontier analysis, we study how deployment decisions may change under possible streamed-token delivery overheads, highlighting that compute-side inference metrics alone can lead to suboptimal placement for latency-sensitive interactive web services.
发表机构
- University of Cyprus(塞浦路斯大学)
- University of Nicosia(尼科西亚大学)
机构由 AI 辅助整理,请以论文原文为准。