arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Hydra:面向边缘SoC代际、后端及量化级别的LLM推理的阶段感知工作负载表征

Hydra: Phase-Aware Workload Characterization of LLM Inference across Edge SoC Generations, Backends, and Quantization Levels

Amir Taherin, Sana Taghipour Anvari, Charles Amante, Yixiao Chen, Ruben Noroian, Zlatan Feric, Nicolas Bohm Agostini, Pu Zhao, José Cano, Bin Ren, Yanzhi Wang, David Kaeli

arXiv 2608.25053首次发表:更新:

发表机构

Northeastern University; University of Glasgow; College of William & Mary(东北大学; 格拉斯哥大学; 威廉玛丽学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

Hydra是面向边缘SoC的LLM推理阶段感知工作负载表征框架,可评估多代SoC、多种LLM及执行格式,揭示延迟、量化、SoC代际对性能的影响,相关成果开源发布。

AI 中文摘要

边缘LLM部署不仅受模型规模和精度影响,推理后端、硬件平台、内存流量及电源管理均会影响延迟与效率。我们提出Hydra,这是一种面向边缘SoC上LLM推理的通用模式、阶段感知工作负载表征框架。Hydra通过共享的逐提示时间模式检测HuggingFace Transformers和vLLM,并将这些记录与硬件遥测数据融合,实现预填充(prefill)和解码(decode)阶段的性能、系统资源利用率及效率的多维表征。利用Hydra,我们评估了三代连续的边缘片上系统(SoC)(AGX Xavier、AGX Orin、AGX Thor)、7个系列的13个指令调优LLM、5种执行格式,并考虑了输入/输出长度敏感性。生成的制品包含约10.7万条逐提示记录,且与Hydra一同公开发布。我们的分析表明,仅总延迟会掩盖关键部署效应:后端结构改变了延迟产生的位置,量化减少了内存流量和能耗但无法单调预测功耗,SoC代际改变了利用率和效率的解读方式。通过将阶段级时间与系统资源利用率及效率指标关联,Hydra可实现可复现的边缘LLM推理的阶段感知表征。Hydra的源代码及收集的逐提示轨迹语料以开源形式提供,网址为:[此处为论文提供的链接]

英文摘要

Edge LLM deployment is shaped by more than model size and precision: inference backend, hardware platform, memory traffic, and power management all affect latency and efficiency. We present Hydra, a common-schema, phase-aware workload characterization framework for LLM inference on edge SoCs. Hydra instruments HuggingFace Transformers and llama.cpp with a shared per-prompt timing schema and fuses those records with hardware telemetry, enabling a multi-dimensional characterization of performance, system-resource utilization, and efficiency across prefill and decode phases. Using Hydra, we evaluate three consecutive edge System-on-Chip (SoC) generations (AGX Xavier, AGX Orin, and AGX Thor), 13 instruction-tuned LLMs from seven families, five execution formats, and consider input/output-length sensitivity. The resulting artifact contains roughly 107K per-prompt records and is publicly released with Hydra. Our analysis shows that aggregate latency alone hides key deployment effects: backend structure changes where latency is introduced, quantization reduces memory traffic and energy but does not predict power monotonically, and SoC generation changes how utilization and efficiency should be interpreted. By connecting phase-level timing with system-resource utilization and efficiency metrics, Hydra enables reproducible, phase-aware characterization of edge LLM inference. Hydra's source code and the collected per-prompt trace corpus are available open-source at: https://github.com/amirtaherin/hydra

CommentsAccepted at the IEEE International Symposium on Workload Characterization (IISWC 2026)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑