AI 中文总结
研究在无服务器Apache Spark环境中阶段级执行器分配问题,提出训练树集成模型估计运行时间和成本,据此为各阶段推荐资源的方法,经实验在TPC-DS和SQLStorm基准测试中实现成本节省和性能平衡。
AI 中文摘要
在分布式处理系统中分配执行器(即计算资源)时,必须在不必要的横向扩展资源成本与人为的性能限制瓶颈之间取得平衡。朴素方法可能在应用程序级别分配执行器,其成本和性能可预测,但对于应用程序执行的数千个不同的单个阶段几乎肯定不是最优的。用户可能还有明确的偏好,如在特定时间预算内完成应用程序并最小化成本,而现有解决方案通常无法支持。我们提出了一种在无服务器Apache Spark环境中确定每个阶段执行器数量的新方法,让用户能够指定所需的成本-性能权衡。我们的方法训练树集成模型来估计阶段的运行时间和成本作为分配资源的函数。然后这些估计用于为每个阶段单独推荐资源。我们在TPC-DS和SQLStorm基准测试中评估了我们的方法,并与两个基线进行了比较。根据用户定义的权衡参数和设置,我们的方法在103个TPC-DS查询中实现了约50%的成本节省,速度仅减慢约16%;在96个SQLStorm查询中实现了约40.5%的成本节省,速度减慢约29%。
英文摘要
Allocating executors (i.e. compute resources) to distributed processing systems must balance resource costs of scaling-out unnecessarily against artificial, performance-limiting bottlenecks. Naive approaches may allocate executors at the application level, which have predictable costs and performance but are almost guaranteed to be sub-optimal for each of the thousands of diverse, individual stages executed by the application. Users may also have explicit preferences, such as completing an application within a specific time budget while minimizing cost, that existing solutions usually fail to support. We propose a novel method for determining the number of executors per stage in a serverless Apache Spark environment, enabling users to specify their desired cost-performance tradeoff. Our approach trains tree-ensemble models to estimate the run times and costs of a stage as a function of allocated resources. These estimates are then used to recommend resources for each stage individually. We evaluate our approach on TPC-DS and SQLStorm benchmarks and compare it against two baselines. Depending on the user-defined trade-off parameter and setup, our approach achieves approx. 50% cost savings across 103 TPC-DS queries with only a approx. 16% slowdown, and approx. 40.5% on 96 SQLStorm queries at a approx. 29% slowdown.