Grouper: 多租户微秒级微服务的调度组
Grouper: Scheduling Groups for Multi-Tenant Microsecond-Scale Microservices
- Sharif University of Technology(谢里夫理工大学)
- Iran University of Science and Technology(伊朗科技大学)
- KTH Royal Institute of Technology(皇家理工学院)
- Amirkabir University of Technology(阿米尔卡比尔理工大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对多租户微秒级微服务,提出调度组机制,通过核心直接传递降低分配器负载,相比现有系统提升中位数和尾部延迟性能。
AI中文摘要:
微秒级核心分配使得将延迟关键型服务与批处理工作负载共置变得有价值。一个线程在微秒内找不到工作就会停车,其核心转给批处理任务。恢复一个核心需要约18微秒,因为分配器必须发现某个核心被需要,然后从持有它的批处理任务中取回。单体架构每个请求支付一次该开销,微服务链在每个跳数(双向)都支付,而多租户主机又将其放大,因为每个租户的跳数在同一分配器处排队。在我们移植的DeathStarBench的hotelReservation上,从两个租户增加到十个租户,单跳延迟从39微秒增至222微秒,10-RPC路径的中位数从456微秒增至2,445微秒,即使没有租户自身负载变化,也造成五倍性能下降。我们提出Grouper和调度组,即一组隔离的运行时,分配器将其视为一个分配和记账单元,其成员可以直接相互传递核心。发送RPC的服务通过无特权内核快速路径将核心捐赠给对端,使核心跟随请求遍历调用图。分配器通过协调、核心寻址撤销和池化预算保留控制权,但离开关键路径;其负载从请求率R和跳数H的Θ(R·H)降至Θ(R)。在每租户每秒1,000-30,000个请求、2至10个租户的网格中,Grouper在中位数上比Caladan(Junction也基于的分配器)和Linux分别快最多7.9倍和3.4倍,在尾部快4.1倍和14.2倍,并在超过70%的负载点上比Caladan为批处理工作留下更多吞吐量。
英文摘要:
Microsecond-scale core allocation makes colocating latency-critical services with batch work worthwhile. A thread that finds no work parks within microseconds and its core goes to a batch task. Putting one back costs $\sim$18 $μ$s, as the allocator must discover that a core is wanted and then take it from the batch task holding it. A monolith pays that tax once per request, a microservice chain pays it at every hop in both directions, and a multi-tenant host multiplies it again, because every tenant's hops queue at the same allocator. On our port of DeathStarBench's hotelReservation, going from two tenants to ten takes a hop from 39 to 222 $μ$s and a 10-RPC path's median from 456 to 2,445 $μ$s, a fivefold degradation even though no tenant's own load changed. We introduce Grouper and the scheduling group, a set of isolated runtimes that the allocator treats as one allocation and accounting unit, whose members may hand cores directly to one another. A service sending an RPC donates its core to the peer through an unprivileged kernel fast path, so the core follows the request through the call graph. The allocator retains control through reconciliation, core-addressed revocation and a pooled budget but leaves the critical path; its load falls from $Θ(R \cdot H)$ to $Θ(R)$ in request rate $R$ and hop count $H$. Over a grid of two to ten tenants at 1,000-30,000 requests per second each, Grouper outperforms Caladan (the allocator Junction also builds on) and Linux by up to 7.9$\times$ and 3.4$\times$ at the median and 4.1$\times$ and 14.2$\times$ at the tail, and leaves batch work more throughput than Caladan at over 70% of load points.