发表机构
Zhejiang University; Alibaba Cloud Computing(浙江大学; 阿里巴巴云计算)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对云原生无服务器数据仓库的资源配置陷阱,提出ScaleSense框架,通过查询编码器、分位数资源预测器与自动扩缩容控制器实现性能-成本优化,在生产查询评估中成效显著。
AI 中文摘要
云原生无服务器数据仓库通过存储与计算解耦实现细粒度弹性扩展,但为高度异构的即席查询确定最优资源分配仍是工业界面临的艰巨挑战。我们对阿里云AnalyticDB生产工作负载的分析揭示了一种代价高昂的“配置陷阱”:用户因担心灾难性资源耗尽而盲目过度配置资源,浪费了巨额资金预算,却未缓解非CPU瓶颈(如I/O饱和)。为突破这一困境,我们提出ScaleSense,一种主动的查询级资源扩缩容框架。该框架包含多维度查询编码器,可联合建模计划拓扑与硬件规格;关键在于,基于分位数的资源预测器能估算多维物理足迹,为最优资源扩缩容提供可靠安全保障;自动扩缩容控制器则可在性能-成本帕累托前沿中导航,根据特定业务优先级动态调整资源分配,无需模型重训练。对超136万条生产查询的评估显示,ScaleSense实现了先进的预测精度与良好的预测区间覆盖率;相较于最优基准,其在最优资源配置选择上取得76.7%的相对提升,解决了关键的性能-成本权衡问题,同时保持低开销推理延迟,验证了其在生产部署中的实际性能;在性能优化策略下,ScaleSense可满足用户定义的性能要求,同时将资金成本降低至多5.22倍。
英文摘要
Cloud-native serverless data warehouses achieve fine-grained elasticity by decoupling storage from compute, yet determining the optimal resource allocation for highly heterogeneous ad-hoc queries remains a formidable industrial challenge. Our analysis of production workloads in Alibaba AnalyticDB exposes a costly ``provisioning trap'': the fear of catastrophic resource depletion drives users to blindly over-provision resources, wasting immense monetary budgets without alleviating non-CPU bottlenecks (e.g., I/O saturation). To break this impasse, we propose ScaleSense, a proactive, query-level resource scaling framework. Specifically, it features a multi-faceted query encoder that jointly models plan topologies and hardware specifications. Crucially, a quantile-based resource predictor estimates multi-dimensional physical footprints, acting as a reliable safety net for optimal resource scaling. An auto-scaling controller then navigates the performance-cost Pareto frontier, dynamically tailoring allocations to specific business priorities without requiring model retraining. Evaluations on over 1.36 million production queries show that ScaleSense achieves state-of-the-art prediction accuracy with good prediction interval coverage. By achieving a 76.7% relative improvement in optimal resource configuration selection over the best baseline, this approach addresses the critical performance-cost trade-off while maintaining low-overhead inference latency, confirming its practical performance in production deployments. Under the performance-optimization policy, ScaleSense satisfies user-defined performance requirements while reducing monetary cost by up to 5.22x.
CommentsThis paper has been accepted for presentation at VLDB 2026