arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30270cs.LG

HybridInfer:面向设备端、边缘和云端LLM推理的热感知强化学习分层路由

HybridInfer: Thermal-Aware Reinforcement-Learning Tier Routing for On-Device, Edge, and Cloud LLM Inference

Simran Koul

首次发表
浏览论文内容

中文总结 AI 辅助

HybridInfer提出热感知强化学习路由器,利用设备端热余量在设备、边缘和云LLM层间路由,解决设备端推理热崩溃问题,在真实Android基准上以最低成本显著提升质量。

中文摘要 AI 辅助

使用小型语言模型的设备端推理可将用户数据保留在本地、离线工作且不产生每次查询的成本,因此当设备端层足够时,它是首选。然而,设备端推理受到热约束的限制,我发现该约束比速度下降更为严重:在旗舰级骁龙设备上,持续的设备端生成会破坏GPU推理运行时的稳定性,在连续几次查询后导致崩溃或静默卡死。故障源于当前工具链(移动GPU上的OpenCL内核编译和长提示预填充),即使设备冷却时也会复发,且对长生成最为严重。跨设备端、边缘和云模型的多层路由器可以缓解此压力,但现有路由器对热状态不敏感,且通常在模拟或非移动硬件上评估。我提出HybridInfer,一种用于三层体系(设备端Llama 3.2 3B、带检索的边缘Llama 3.1 8B、云端GPT-4o)的热感知强化学习路由器,它使用手机的热余量和查询复杂度估计作为状态,并通过离线训练的Q学习策略选择层。其奖励在质量与延迟、成本和热惩罚之间权衡,外加奖励设备端执行的本地性奖励。我证明该奖励是热感知路由的先决条件:没有它,最优策略会卸载所有查询。在210个提示的真实Android基准上,学习到的路由器在最低成本下显著优于两个手动调整的启发式方法(配对Wilcoxon检验,p < 0.02)。始终设备端条件在可服务查询上匹配每次查询质量,但慢三到六倍,且在长查询上失败,因此路由在延迟、可靠性和覆盖范围上获胜,而非质量。据我所知,这是首次在真实硬件上利用设备端热余量来选择不同能力的LLM推理层。

英文摘要

On-device inference with small language models keeps user data local, works offline, and incurs no per-query cost, so the on-device tier is preferred when it is adequate. It is thermally constrained, however, and I find the constraint is sharper than a slowdown: on a flagship Snapdragon device, sustained on-device generation destabilizes the GPU inference runtime, which crashes or silently wedges after a few consecutive queries. The failure lies in the current toolchain (OpenCL kernel compilation and long-prompt prefill on the mobile GPU), recurs even when the device is cool, and is worst for long generations. Multi-tier routers across on-device, edge, and cloud models can relieve this pressure, but existing routers are thermal-blind and typically evaluated in simulation or on non-mobile hardware. I present HybridInfer, a thermal-aware reinforcement-learning router for a three-tier hierarchy (on-device Llama 3.2 3B, edge Llama 3.1 8B with retrieval, cloud GPT-4o) that uses the phone's thermal headroom and a query-complexity estimate as state and selects a tier by an offline-trained Q-learning policy. Its reward trades quality against latency, cost, and a thermal penalty, plus a locality bonus crediting on-device execution. I show this bonus is a precondition for thermal-aware routing: without it the optimal policy offloads every query. On a real Android benchmark of 210 prompts, the learned router attains significantly higher quality than two hand-tuned heuristics (paired Wilcoxon, p < 0.02) at the lowest cost of any adaptive condition. Always-on-device conditions match per-query quality on servable queries but are three to six times slower and fail on long queries, so routing wins on latency, reliability, and coverage rather than quality. To my knowledge this is the first use of on-device thermal headroom to select among LLM inference tiers of differing capability on real hardware.

补充信息

↑