arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.00267eess.SYcs.SY

ThermE:热受限边缘SoC上持续LLM推理的共享热余量预测管理

ThermE: Predictive Management of Shared Thermal Headroom for Sustained LLM Inference on Thermally Constrained Edge SoCs

  • The Hong Kong Polytechnic University(香港理工大学)
  • Southern University of Science and Technology(南方科技大学)

机构由 AI 辅助整理,请以论文原文为准。

Rui Lu, Yuheng Wang, Bozheng Liu

AI总结:

针对边缘SoC上持续LLM推理的热约束问题,提出ThermE运行时系统,通过快速热编译、PDE约束预测和不确定性感知调度联合管理共享热余量,在vLLM上实现TTFT和TPOT分别降低40.55%和12.17%,SLO违反率降至5.70%。

AI中文摘要:

紧凑型边缘片上系统(SoC)平台在热约束下越来越多地运行持续的大语言模型(LLM)推理,而其CPU、GPU和RAM共享同一条冷却路径。因此,预填充(prefill)和解码(decode)阶段消耗共享且随时间变化的热余量,然而厂商自带的调控器仅在接近硬件降频阈值时才做出反应,且不了解请求状态或即将到来的工作负载。一个能提升当前性能的控制决策可能会过快消耗热余量,从而降低后续服务质量。本文提出ThermE,一个运行时系统,用于预测并联合管理边缘SoC上持续LLM推理的共享热余量。其快速LLM到热编译器(Fast LLM-to-Heat Compiler)在不执行LLM的情况下,将模型、请求、运行时状态和候选动作映射到领域热量。基于偏微分方程(PDE)约束的热余量预测器(Headroom Predictor)使用ThermPINN进行离线热辨识,并使用降阶热余量预测器(RHP)进行低开销的在线不确定性校准热余量预测。随后,一个不确定性感知的动作调度器(Uncertainty-Aware Action Scheduler)选择能在服务质量和未来热余量之间取得平衡的动作。我们在vLLM之上实现ThermE,并在四个LLM推理工作负载上进行评估。结果表明,相对于vLLM,ThermE将TTFT和TPOT分别降低了40.55%和12.17%。其SLO违反率为5.70%,而最强基线为12.30%;其预测器取得了1.94°C的MAE,开销为11.50毫秒。

英文摘要:

Compact edge system-on-chip (SoC) platforms increasingly run sustained LLM inference under thermal constraints, while their CPU, GPU, and RAM share a cooling path. Prefill and decode therefore consume shared, time-varying thermal headroom, yet vendor governors react only near hardware throttling thresholds without knowledge of request state or upcoming work. A control decision that improves current performance can thus consume headroom too quickly and degrade subsequent service. In this paper, we present ThermE, a runtime system that predicts and jointly manages shared thermal headroom for sustained LLM inference on edge SoCs. Its Fast LLM-to-Heat Compiler maps the model, requests, runtime state, and candidate actions to domain heats without executing LLMs. A partial differential equation (PDE)-constrained Headroom Predictor uses ThermPINN for offline thermal identification and a Reduced Headroom Predictor (RHP) for low-overhead online uncertainty-calibrated headroom forecasts. An Uncertainty-Aware Action Scheduler then selects actions that balance serving quality and future headroom. We implement ThermE atop vLLM and evaluate it across four LLM inference workloads. The results show that ThermE reduces TTFT and TPOT by 40.55% and 12.17%, respectively, relative to vLLM. It achieves a 5.70% SLO violation rate, compared with 12.30% for the strongest baseline, while its predictor obtains a 1.94 $^\circ$C MAE with 11.50 ms overhead.

补充信息

↑