arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.11770cs.DB

AutoSLO:云数据仓库上的实用延迟服务水平目标——扩展版本

AutoSLO: Practical Latency SLOs on Cloud Data Warehouses -- Extended Version

Markos Markakis, Tim Kraska

AI总结:

研究云数据仓库中多集群工作负载管理问题,提出AutoSLO框架,通过策略调优器、自动缩放器和查询路由器三个组件在三个时间尺度运行,能满足不同延迟SLO,降低成本,各组件有效降低SLO违反率。

AI中文摘要:

现代云数据仓库将计算与存储解耦,便于组织使用多个计算集群访问相同基础数据。这种灵活性常用于不同工作负载间的性能隔离,使各工作负载更可靠地满足延迟服务水平目标(SLO)。但专用集群方法需持续扩展计算集群以适应工作负载变化,过度配置浪费资源,配置不足则有违反SLO风险。我们提出AutoSLO,一个用于多集群云数据仓库的延迟SLO感知工作负载管理框架。AutoSLO通过三个关键组件在三个时间尺度上运行。首先,周期性策略调优器利用历史工作负载预测模拟规划主动集群扩展行动并调整配置参数。其次,SLO感知反应式自动缩放器在近期工作负载行为偏离预测时调整活动集群集。第三,在线查询路由器在放置每个查询时对实时负载做出反应,使用并发感知延迟预测器避免违反SLO。在实际Redbench工作负载上,AutoSLO成功满足不同严格程度的延迟SLO,与每个场景下的次优基线相比,平均成本降低26.4%。组件级评估表明,查询路由器和自动缩放器相对于相应替代方案分别将SLO违反率平均降低47.8%和93.7%。最后,我们表明策略调优器使用一天的工作负载历史记录可将SLO违反率平均降低44.6%,且每个组件在其预期运行时间尺度下都很高效。

英文摘要:

Modern cloud data warehouses decouple compute from storage, making it easy for organizations to access the same underlying data with multiple compute clusters. This flexibility is often used for performance isolation among diverse workloads, so that each workload meets its latency service-level objective (SLO) more reliably. For example, interactive dashboards, ad hoc analysis, and batch jobs can each run on separate clusters. However, this dedicated-cluster approach requires each compute cluster to be continuously scaled to adapt to workload evolution, with over-provisioning wasting resources and under-provisioning risking SLO violations. We present AutoSLO, a latency-SLO-aware workload management framework for multi-cluster cloud data warehouses. AutoSLO operates across three timescales through three key components. First, a periodic Policy Tuner plans proactive cluster scaling actions and tunes configuration parameters, using simulations of history-derived workload forecasts. Second, an SLO-aware reactive Autoscaler adjusts the active cluster set when recent workload behavior deviates from the forecast. Third, an online Query Router reacts to live load when placing each query, using a concurrency-aware latency predictor to avoid SLO violations. On realistic Redbench workloads, AutoSLO successfully meets latency SLOs of varying strictness, reducing cost by a mean of 26.4% compared to the per-scenario next-best baseline. Component-level evaluations show that the Query Router and Autoscaler respectively reduce SLO violation rates by a mean of 47.8% and 93.7%, relative to their corresponding alternatives. Finally, we show that the Policy Tuner can reduce the SLO violation rate by a mean of 44.6% using a single day of workload history, and that each component is efficient given its intended operating timescale.

补充信息

↑