AI 中文总结
提出开源Kubernetes调度框架插件OPSche,结合外部约束优化求解器,通过三种触发模式及阻塞变体优化集群部署,提升资源使用率、降低调度延迟。
AI 中文摘要
Kubernetes是当前最先进的容器编排工具,其默认调度器采用快速的局部部署决策,但该设计会导致资源碎片化、集群使用率降低以及过度配置。外部求解器可计算全局部署计划,但在上游集群中执行这些计划难度较大:Kubernetes无原生跨节点抢占机制,未协调的并发调度会导致不一致,替换默认调度器会使部署脱离上游周期。我们提出OPSche,这是一款开源的Kubernetes调度框架插件,外部求解器可结合默认调度器驱动集群范围的部署决策。OPSche通过协调的框架钩子原子性地验证并执行求解器生成的计划,支持三种触发模式:调度失败模式(工作负载无法部署时触发)、周期性模式(固定时间间隔触发)、稳定队列模式(待处理工作负载集稳定时触发),每种模式均有阻塞变体,可更精细地调整部署质量、延迟和中断情况。我们将OPSche与基于约束的优化求解器结合,在广泛的集群配置中验证了其可行性,结果显示资源使用率提升最高达3.0%,调度延迟降低超过1秒。
英文摘要
The default scheduler of Kubernetes, the state-of-the-art container orchestrator, uses fast, local placement decisions. Unfortunately, this design leads to resource fragmentation, reduced cluster usage, and overprovisioning. External solvers can compute global placement plans, but enforcing these plans in upstream clusters is hard. Kubernetes provides no native cross-node preemption, uncoordinated concurrent scheduling leads to inconsistencies, and replacing the default scheduler would sever deployments from upstream cycles. We present OPSche, an open-source Kubernetes Scheduling Framework plugin where external solvers can drive cluster-wide placement decisions in concert with the default scheduler. OPSche atomically validates and enforces solver-produced plans through coordinated framework hooks and supports three trigger modes: scheduling-failure, periodic, and stable-queue -- resp. triggered when a workload cannot be placed, at fixed time intervals, when the set of pending workloads stabilises. Each mode has a blocking variant for a finer tuning of placement quality, latency, and disruption. We pair OPSche with a constraint-based optimisation solver, showing its feasibility across a broad set of cluster configurations and reporting improvements of resource usage by up to 3.0% and scheduling latency by more than a second.