arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.18390cs.DC

SLO-Scaler:面向微服务的感知不确定性、由服务水平目标(SLO)驱动的自动扩缩容方案

SLO-Scaler: Uncertainty-Aware SLO-Driven Autoscaling for Microservices

Shuo Wang, Xiaoxuan Sun, Shao-yu Huang, Bencheng Su, Shuo Xu, Netra Awate

AI总结:

SLO-Scaler 是一款感知不确定性的微服务自动扩缩容框架,采用贝叶斯 LSTM 模型预测关键指标,结合依赖图分析定位瓶颈,在突发流量下可显著降低 SLO 违反率、减少资源消耗与扩缩容频率并优化延迟。

AI中文摘要:

自动扩缩容基于微服务的应用以满足服务水平目标(SLO)仍面临诸多挑战,包括突发 workload、服务依赖间的级联延迟以及冷启动开销。现有方法如 Kubernetes 水平 Pod 自动扩缩器(HPA)依赖基于阈值的 CPU 或内存指标,对流量峰值反应过慢。近期的预测方法提升了响应速度,但生成的点预测忽略了预测不确定性,导致过度配置或扩缩容振荡。我们提出 SLO-Scaler,这是一个感知不确定性的自动扩缩容框架,采用贝叶斯 LSTM 模型预测短时间范围内的请求速率、尾部延迟和 SLO 违反概率。SLO-Scaler 将基于置信区间的扩缩容决策与定位瓶颈服务的依赖图分析模块相结合,避免不必要的全链路扩缩容。我们在部署于 Kubernetes 上的 DeathStarBench 社交网络基准测试中,针对周期性、突发和长尾流量模式评估了 SLO-Scaler。在突发流量下,与基线方法相比,SLO-Scaler 将 SLO 违反率降低了 29%-56%,平均副本数减少了 18%-33%,扩缩容事件频率降低了 38%-59%,同时实现了更低的尾部延迟。

英文摘要:

Autoscaling microservice-based applications to satisfy Service Level Objectives (SLOs) remains challenging due to bursty workloads, cascading latency across service dependencies, and cold-start overhead. Existing approaches such as the Kubernetes Horizontal Pod Autoscaler (HPA) rely on threshold-based CPU or memory metrics, which react too slowly to traffic spikes. Recent predictive methods improve responsiveness but generate point forecasts that ignore prediction uncertainty, leading to over-provisioning or oscillatory scaling. We propose SLO-Scaler, an uncertainty-aware autoscaling framework that predicts short-horizon request rates, tail latency, and SLO violation probability using a Bayesian LSTM model. SLO-Scaler integrates confidence-interval-based scaling decisions with a dependency graph analysis module that localizes bottleneck services, avoiding unnecessary whole-chain scaling. We evaluate SLO-Scaler on the DeathStarBench Social Network benchmark deployed on Kubernetes under periodic, bursty, and long-tail traffic patterns. Under bursty traffic, SLO-Scaler reduces the SLO violation rate by 29-56%, lowers the average replica count by 18-33%, and decreases scaling event frequency by 38-59% compared with the baselines, while achieving lower tail latency.

↑