AI 中文总结
HW-Router是一种动态路由框架,结合实时硬件信号与模型特征实现SLO感知路由,在多LLM服务中相比CARROT等基线显著降低延迟、提升SLO达成率并均衡GPU负载。
AI 中文摘要
现代大语言模型(LLM)服务平台会在不同GPU上部署多个模型,需要路由将传入的查询定向到合适的LLM。然而,现有的路由方法主要依赖静态模型属性(如规模或浮点运算量FLOPs)来估算服务成本。这种静态成本建模无法捕捉实际部署中的动态行为,同一模型的推理延迟会因硬件类型(如H100与V100)、当前系统负载(如运行队列和等待队列长度)以及资源竞争(如KV缓存使用率和GPU利用率)而表现出巨大差异。这种与硬件无关的路由会导致决策不当,引发服务水平目标(SLO)违规、队列堆积和GPU利用率不足等问题。为应对这些挑战,本文提出HW-Router,这是一种动态路由框架,它将实时硬件信号融入模型选择,以实现准确的延迟预测和智能的、SLO感知的路由决策。该方法结合了模型特定特征(架构、规模、输入长度)与硬件指标,包括队列长度、KV缓存利用率以及最近的首字符生成时间(TTFT)/每输出令牌时间(TPOT)性能,并使用轻量级延迟预测器估算每个模型在每个GPU上的服务时间。在各类工作负载上的评估显示,与最先进的路由基线CARROT和IRT相比,HW-Router实现了3.4-3.9倍的端到端延迟降低、46-48个百分点的SLO达成率提升、6-8倍的GPU负载偏差降低,以及等待队列占比降低3.1-3.4倍,且仅产生约200微秒的额外路由开销,同时未出现输出质量损失。这些结果凸显了实时硬件反馈对于可扩展、可预测且均衡的多LLM服务的重要性。代码可在指定URL获取。
英文摘要
Modern large language model (LLM) serving platforms deploy multiple models across different GPUs, requiring routers to direct incoming queries to appropriate LLMs. However, existing routing approaches primarily rely on static model attributes such as size or FLOPs to estimate serving costs. This static cost modeling fails to capture the dynamic behavior of real deployments, where the same model can exhibit vastly different inference latencies depending on hardware type (e.g., H100 vs. V100), current system load (e.g., running and waiting queue lengths), and resource contention (e.g., KV-cache usage and GPU utilization). Such hardware-agnostic routing leads to suboptimal decisions, resulting in SLO violations, queue buildup, and underutilized GPUs. To address these challenges, we present HW-Router, a dynamic routing framework that integrates real-time hardware signals into model selection to enable accurate latency prediction and intelligent, SLO-aware routing decisions. Our approach incorporates model-specific features (architecture, size, input length) alongside hardware metrics including queue lengths, KV-cache utilization, and recent TTFT/TPOT performance, and uses a lightweight latency predictor to estimate per-model-per-GPU serving time. Evaluations across diverse workloads show that HW-Router achieves 3.4-3.9x lower end-to-end latency, 46-48 percentage points higher SLO attainment, 6-8x lower GPU load skew, and a 3.1-3.4x reduction in waiting-queue fraction compared to state-of-the-art router baselines, CARROT and IRT, with only ~200 us of additional routing overhead and no loss in output quality. These results highlight the importance of real-time hardware feedback for scalable, predictable, and well-balanced multi-LLM serving. Code is available at https://github.com/UCF-ML-Research/HW-Router.
CommentsPreprint
Journal ref2026 Design Automation Conference (DAC)