arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.29103cs.DC

CLASP:面向无服务器流处理的链式请求感知扩缩容与算子部署

CLASP: Chained-Request-Aware Scaling and Operator Placement for Serverless Stream Processing

  • The University of Melbourne(墨尔本大学)

机构由 AI 辅助整理,请以论文原文为准。

Tianyu Qi, Maria A. Rodriguez, Rajkumar Buyya

AI总结:

针对无服务器流处理中链式请求开销被忽视导致的工作节点数量估计偏差问题,提出CLASP策略,通过联合考虑执行与链式请求开销优化算子部署,实现吞吐量提升与延迟降低。

AI中文摘要:

带状态的无服务器(函数即服务,Function-as-a-Service)环境中,其工作节点托管状态服务器,正越来越多地被用于流处理。流应用是由算子组成的流水线,每个算子通过链式请求向下游转发中间数据。当输入速率波动时,系统应调整算子并行度并在工作节点间部署实例,以维持输入速率。现有方法在执行时未充分考虑链式请求的开销,导致其对所需工作节点数量的估计出现偏差:工作节点过少会使集群无法跟上输入速率,过多则会导致更大比例的链式请求跨工作节点传输,从而增加端到端延迟。我们提出CLASP,一种面向带状态无服务器环境下流处理的扩缩容与调度策略。在运行时,CLASP从观测指标中估计执行开销和链式请求开销,在覆盖这两种开销的容量模型下,调整算子并行度并将算子打包到能维持目标输入速率的最少工作节点上。一旦做出扩缩容决策,CLASP会将每个算子的状态与其实例一同迁移,从而最小化执行暂停时间。实验表明,与最先进的扩缩容策略相比,CLASP的吞吐量最高提升3.3倍,端到端中位数延迟最高降低76%。

英文摘要:

Stateful serverless (Function-as-a-Service) environments, whose workers host state servers, are increasingly used for stream processing. A stream application is a pipeline of operators, where each operator forwards intermediate data downstream through a chained request. As input rates fluctuate, the system should adjust operator parallelism and place instances across workers to sustain the incoming rate. Existing approaches do so without fully accounting for chained-request overhead, leading them to misestimate the required number of workers. Too few leave the cluster unable to keep up with the input rate, while too many route a larger fraction of chained requests across worker boundaries, increasing end-to-end latency. We propose CLASP, a scaling and scheduling strategy for stream processing in stateful serverless environments. At runtime, CLASP estimates execution cost and chained-request cost from observed metrics. Under a capacity model that covers the two costs, it adjusts operator parallelism and packs operators onto the fewest workers that can sustain the target input rate. Once a scaling decision is made, CLASP migrates each operator's state together with its instances, thereby minimizing execution pause time. Experiments show that CLASP improves throughput by up to 3.3x and reduces median end-to-end latency by up to 76% compared with state-of-the-art scaling strategies.

↑