发表机构
DiDi Chuxing(滴滴出行)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SchedBlame是基于标准内核的eBPF跟踪工具,可将容器CPU竞争归因于罪魁祸首控制组,能拆分容器CPU需求、标记异常,在生产环境中开销较低。
AI 中文摘要
共享同一台机器的容器会争夺CPU资源,当其中一个容器运行变慢时,运维人员需要知道是哪个同租户容器导致的,但当前已部署的信号都无法给出答案。压力停滞信息、每个控制组(cgroup)的等待计数器以及运行队列延迟直方图都是受害者侧的工具:它们仅报告某个容器等待了多久,却从不说明它在等待谁。要找出罪魁祸首,需要对内核打补丁、进行完整的调度器跟踪或采用统计推断方法:这些方法要么不可移植,要么成本过高无法持续启用,要么在多个受害者共存时不可靠。SchedBlame是一种eBPF跟踪工具,可在标准内核上持续将CPU竞争归因于引发竞争的控制组。它反转了核算逻辑:不再测量受害者的等待时长,而是测量在该受害者处于可运行状态但未在同一CPU上运行时,其他每个控制组消耗的CPU时间。其机制是一个每个CPU的位图,用于记录哪些被测控制组处于等待状态,该位图通过内核在四个调度器钩子处的自身可运行计数进行维护。每个运行切片都携带该位图,因此一条16字节的记录可将CPU时间计入完整的竞争者-受害者归因矩阵的一行;内核不存储任何成对的状态。由此产生三个特性:切片是自描述的,因此用户空间不存储等待状态,丢失一条记录只会影响测量结果,不会影响正确性;被测集合可通过发布一个周期(epoch)进行重新配置,在钩子持续运行的同时,以恒定时间使所有缓存和每个CPU位图失效;采样从不触及等待状态,因此通过逆保留概率进行重新缩放可保持估计量无偏。SchedBlame将每个容器每秒的CPU需求拆分为运行时间、内部竞争、外部竞争和节流时间,针对滚动的第99百分位基准标记异常,并命名相关的竞争容器。在未修改的4.18和5.10内核的生产环境中,该工具在96核主机上跟踪84个容器,其开销约为Redis吞吐量的1%和一个核心的6%。
英文摘要
Containers that share a machine compete for CPU. When one slows down, the operator needs to know which co-tenant is responsible, and no deployed signal can say. Pressure stall information, per-cgroup wait counters, and run-queue latency histograms are all victim-side: they report that a container waited, never who it waited for. Recovering the culprit means a kernel patch, full scheduler tracing, or statistical inference: unportable, too costly to leave on, or unreliable when victims coexist. SchedBlame is an eBPF tracer that attributes CPU contention to the cgroups that caused it, on stock kernels, continuously. It inverts the accounting: instead of measuring how long a victim waited, it measures the CPU time every other cgroup consumed while that victim was runnable but not running on the same CPU. The mechanism is a per-CPU bitmap of which measured cgroups are waiting, maintained from the kernel's own runnable counts at four scheduler hooks. Every run slice carries that bitmap, so one 16-byte record charges CPU time to a full row of a competitor x victim blame matrix; the kernel stores no per-pair state. Three properties follow. Slices are self-describing, so userspace holds no waiting state and a lost record costs measurements, not correctness. The measured set is reconfigured by publishing an epoch, invalidating every cache and per-CPU bitmap in constant time while the hooks keep running. Sampling never touches waiting state, so rescaling by the inverse keep probability keeps the estimator unbiased. SchedBlame splits each container's per-second CPU demand into runtime, internal contention, external contention, and throttling, flags anomalies against a rolling 99th-percentile baseline, and names the competitors responsible. In production on unmodified 4.18 and 5.10 kernels, tracking 84 containers on a 96-core host, it costs about 1% of Redis throughput and 6% of one core.
Comments12 pages, 2 figures. Preliminary evaluation; a full evaluation plan is stated in the paper