arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于受损资源量指标的NoSQL工作负载的以资源为中心的分析与优化

A Resource-centric Analysis and Optimization of NoSQL Workloads using Distressed Resource Volume Metric

Gunika Verma, Aashutosh A, Pooja Srinivas, Yogesh Simmhan, Ayush Choure, Harshit Shah, Mayukh Das, Prashant Sasatte, Chetan Bansal, Abhijit Pai, Suraj Dixit, Achint Agrawal

arXiv 2608.09173首次发表:更新:

AI 中文总结

针对公开NoSQL工作负载缺失的问题,研究人员以Cosmos DB为背景,提出DRV指标、LoadStar框架、Luna模型及Orbit算法,实现资源优化,Orbit可降35%资源、减错误率,已部署生产并年省数亿美元。

AI 中文摘要

大规模托管云数据库利用复杂的负载打包与迁移(PAM)算法,这些算法提供了在云资源上大规模运行这些服务所需的效率。由于缺乏公开的NoSQL工作负载,针对大规模云数据库的资源与可靠性优化研究受到限制。我们在微软旗舰云托管NoSQL数据库Cosmos DB的背景下解决这一问题。我们首先从真实的Cosmos DB集群中提出开源NoSQL工作负载,并分析这些轨迹以推导一种新的可靠性指标——受损资源量(DRV),该指标捕捉终端用户体验到的服务质量。随后,我们开发了一个开源策略模拟框架LoadStar,该框架由用于估计真实流量模式服务质量的非参数统计模型驱动,这些构成了用于验证以资源为中心的NoSQL工作负载策略的可复用基准流水线。接着,我们定义了将Cosmos DB副本部署到VM节点上的资源优化问题,开发了用于预测未来负载分布的Luna模型,以及使用这些预测来触发和重新平衡受压力副本以减少尾部错误的Orbit PAM算法。我们使用LoadStar对这些工作负载进行的实验验证了Orbit相较于现有Cosmos DB策略和最差适配优化基线的优势:在更低错误率下提供更高负载,且资源减少高达35%。这些成果已部署到生产环境,每年潜在节省数亿美元,同时为数百万客户提升服务可靠性。

英文摘要

Large-scale managed cloud databases leverage sophisticated load Packing and Migration (PAM) algorithms, which provide the efficiencies necessary for running these services at scale on cloud resources. Research into optimizing the resources and reliability of cloud databases at massive scales is limited by a lack of public NoSQL workloads. We address this in the context of Cosmos DB, Microsoft's flagship cloud-hosted NoSQL database. We first propose open-source NoSQL workloads from real Cosmos DB clusters, and analyze these traces to derive a novel reliability metric, Distressed Resource Volume (DRV), which captures the quality of service experienced by the end user. We then develop an open-source policy simulation framework, LoadStar, powered by a non-parametric statistical model of estimating the QoS of real traffic patterns. These form a reusable benchmark pipeline for validating policies for resource-centric NoSQL workloads. We then define a resource optimization problem for placing Cosmos DB replicas onto VM nodes, develop the Luna model for forecasting future load distributions, and the Orbit PAM algorithm that uses these forecasts to trigger and rebalance stressed replicas, to reduce tail-errors. Our experiments, validated using LoadStar for these workloads, demonstrate Orbit's benefits over the existing Cosmos DB policy and a worst-fit optimized baseline, with higher load delivered at lower error rates and up to $35\%$ reduction in resources. These have been deployed in production, with potential savings of $\$100M$s/yr while improving service reliability for millions of customers.

CommentsVLDB 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑