面向计算连续体的质量感知推理服务的 Kubernetes 原生请求路由器
A Kubernetes-Native Request Router for Quality-Aware Inference Serving in the Computing Continuum
浏览论文内容
中文总结 AI 辅助
提出 Kubernetes 原生请求路由器 ASRB,基于分数动态路由,联合基础设施与质量指标,平衡延迟与准确性,降低响应时间、故障率和监控成本。
中文摘要 AI 辅助
我们提出了自适应基于分数的路由均衡器(ASRB),这是一种动态的、基于分数的请求路由机制,用于在计算连续体上基于 Kubernetes 的服务部署。ASRB 联合考虑基础设施层面的信息、响应时间测量和应用层面的质量指标,特别关注机器学习(ML)工作负载的推理服务。对于这些工作负载,ASRB 在连续体中部署的服务实例之间平衡请求,遵循服务提供商定义的策略,这些策略编码为 QoS 标准的加权组合,以灵活地处理延迟与准确性之间的权衡。为了驱动路由决策并快速适应运行环境的变化,ASRB 跨多个系统层监控一系列运行时指标。为了处理相关的监控开销(对于大规模部署尤为重要),它选择性地、自适应地控制监控强度,而不牺牲路由质量。ASRB 的实现无需对 Kubernetes 进行任何修改,使其易于在现有集群环境中部署和操作。我们的测试平台实验证明了 ASRB 的多功能性:当针对延迟降低进行调优时,与面向延迟的最先进路由机制相比,它实现了至少 10 毫秒更低的平均响应时间,而当通过特定配置优先考虑准确性时,它实现了更高的准确性,从而实现了灵活且操作者可控制的权衡。同时,它实现了更低的故障率、对运行环境变化的更高响应性,以及比相关最先进解决方案高达约 70% 的监控成本降低,在某些配置中可能仅以适度的延迟惩罚为代价。
英文摘要
We introduce Adaptive Score-based Routing Balancer (ASRB), a dynamic, score-based request routing mechanism for Kubernetes-based service deployments over the computing continuum. ASRB jointly considers infrastructure-level information, response time measurements, and application-level quality indicators, with a particular focus on serving Machine Learning (ML) workloads. For these workloads, ASRB balances requests over service instances deployed in the continuum, following service provider-defined policies encoded as weighted combinations of QoS criteria to flexibly address latency-accuracy trade-offs. To drive routing decisions and swiftly adapt to changes in the operating environment, ASRB monitors a range of runtime metrics across multiple system layers. To deal with the associated monitoring overhead, particularly important for large-scale deployments, it selectively and adaptively controls monitoring intensity without sacrificing on routing quality. ASRB is implemented without requiring any modifications to Kubernetes, making it straightforward to deploy and operate in existing cluster environments. Our testbed experiments demonstrate the versatility of ASRB: When tuned for latency reduction, it achieves at least 10 ms lower mean response time compared with latency-oriented state-of-the-art routing mechanisms, while it achieves higher accuracy when this is prioritized through specific configurations, thus enabling flexible and operator-controllable trade-offs. At the same time, it attains reduced failure rates, higher responsiveness to changes in the operating environment, and up to ~70% less monitoring cost than relevant state-of-the-art solutions, at the potential expense of only a modest latency penalty in some configurations.
发表机构
- Distributed Systems Group, TU Wien(维也纳工业大学分布式系统研究组)
- TU Wien(维也纳工业大学)
- Faculty of Electrical Engineering and Computing, University of Zagreb(萨格勒布大学电气工程和计算学院)
机构由 AI 辅助整理,请以论文原文为准。