AI 中文总结
研究在Kubernetes微服务上维护SLO的难题,提出ARBITER受保护控制平面,构建因果资源图等,分离规划与执行,支持多种规划方式。通过实验评估其在选择修复操作和下游目标等方面的效果,展示了优于HPA/仅资源控制的性能。
AI 中文摘要
在Kubernetes微服务上维护服务水平目标(SLO)仍然很困难,因为自动缩放器观察粗略的资源指标,近期的SLO控制器通常依赖自定义遥测,且无约束的代理操作符无法安全地变更生产集群。我们提出了ARBITER,一个面向SLO的Kubernetes修复的受保护控制平面。ARBITER构建了一个原生OpenTelemetry因果资源图,组装有界的诊断上下文对象,并公开一个有限的类型化操作接口,将规划与执行分离。该接口支持确定性规划器和基于大语言模型的规划工具,通过确定性模式检查、策略门限、资源/中断预算、审批和有界执行形成安全基础。我们在一个4节点的Kubernetes集群上使用DeathStarBench社交网络和在线精品店对ARBITER进行评估。评估测试了仅资源自动缩放无法提供的两种面向SLO的控制形式:选择正确的修复操作和选择正确的下游目标。对于坏镜像部署回归,ARBITER在所有十次CPU消耗和纯延迟运行中都选择回滚金丝雀;水平 Pod 自动缩放器(HPA)要么扩展有故障的镜像,要么从不触发。对于下游关键路径故障,用户可见的故障出现在前端,但追踪证据将主时间线服务识别为可修复的瓶颈。确定性的ARBITER和实时审批门控中的十四行诗工具在每个副本中都针对该下游服务,而HPA/仅资源控制则从未这样做。额外的实验涵盖了受保护的放置修复、在线精品店的可移植性、对抗性安全拒绝、离线多模型重放以及基于KWOK的控制平面扩展证据。我们发布了控制器、重放语料库、工具、安全测试和图形工件:此https URL。
英文摘要
Maintaining service-level objectives (SLOs) on Kubernetes microservices remains difficult because autoscalers observe coarse resource metrics, recent SLO controllers often depend on custom telemetry, and unconstrained agentic operators cannot safely mutate production clusters. We present ARBITER, a guarded control plane for SLO-oriented Kubernetes remediation. ARBITER builds an OpenTelemetry-native causal resource graph, assembles bounded DiagnosisContext objects, and exposes a finite typed-action interface that separates planning from execution. The same interface supports deterministic planners and an LLM-backed planning harness, with deterministic schema checks, policy gates, resource/disruption budgets, approval, and bounded execution forming the safety substrate. We evaluate ARBITER on a 4-node Kubernetes cluster using DeathStarBench Social Network and Online Boutique. The evaluation tests two forms of SLO-oriented control that resource autoscaling alone does not provide: selecting the right remediation action and selecting the right downstream target. For bad-image deployment regressions, ARBITER selects rollback_canary in all ten CPU-burn and pure-latency runs; HPA either scales the faulty image or never triggers. For a downstream critical-path fault, the user-visible breach appears at the frontend, but trace evidence identifies home-timeline-service as the remediable bottleneck. Deterministic ARBITER and a live approval-gated Sonnet harness target that downstream service in every replicate, whereas HPA/resource-only control never does. Additional experiments cover guarded placement repair, Online Boutique portability, adversarial safety rejection, offline multi-model replay, and KWOK-based control-plane scale evidence. We release the controller, replay corpus, harnesses, safety tests, and figure artifacts: https://github.com/pooyan/arbiter.