arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.26503cs.DC

安全门控自动扩缩容:Kubernetes垂直资源优化的多层防御架构

Safety-Gated Autoscaling: A Multi-Layered Defense Architecture for Kubernetes Vertical Resource Optimization

Azra Karakaya, Erva Şengül, Ahmet Kaplan

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对Kubernetes资源浪费与被动扩缩器缺陷,提出带五层安全流水线的开源智能集群优化器,经测试和GKE部署验证,可实现20%-40%的成本节约且内存泄漏检测准确率达83%。

中文摘要 AI 辅助

Kubernetes是编排容器化应用的标准平台,但资源管理仍存在难点。为保障安全,工程师会过度配置CPU和内存,导致预留但未使用的容量成为成本浪费的主要来源。内置的水平和垂直Pod自动扩缩器(Horizontal and Vertical Pod Autoscalers)是被动式的:仅在阈值被突破后才会采取行动,这会导致滞后、过度配置,还可能通过给内存泄漏的工作负载分配更多内存来掩盖软件缺陷。预测式自动扩缩器专注于提升预测准确性或在专有基础设施内运行,异常检测仅用于告警,从不用于阻止有害操作。智能集群优化器(Intelligent Cluster Optimizer)是一款开源Kubernetes算子,它以安全为首要考量对容器工作负载进行合理规格调整。其核心贡献是一个五层安全流水线,其中基于R²评分线性回归的内存泄漏检测器充当阻塞门:若检测到泄漏则拒绝该建议,因此优化器绝不会通过扩大故障容器来掩盖bug。该流水线结合了SLA监控、断路器、HPA/PDB冲突检测和策略引擎,具备回滚和干运行模式供人工审批。建议通过百分位分析和Holt-Winters预测生成,在每个容器层面通过多目标帕累托优化进行平衡。我们用1118个自动化测试(覆盖率达80.3%)和在Google Kubernetes Engine上的实际部署验证了该系统,合理规格调整在假设场景中预计节省20%-40%的成本,泄漏门的检测准确率达83%。

英文摘要

Kubernetes is the standard platform for orchestrating containerized applications, yet resource management remains difficult. To stay safe, engineers over-provision CPU and memory, leaving reserved but unused capacity that is the main source of wasted cost. The built-in Horizontal and Vertical Pod Autoscalers are reactive: they act only after a threshold is crossed, which causes lag, over-provisioning, and can mask software defects by granting a leaking workload more memory. Predictive autoscalers focus on improving forecasting accuracy or run inside proprietary infrastructure, and anomaly detection is used only to alert, never to block a harmful action. The Intelligent Cluster Optimizer is an open-source Kubernetes operator that right-sizes container workloads with safety as a first-class concern. Its central contribution is a five-layer safety pipeline where a memory-leak detector, based on linear regression with R^2 scoring, acts as a blocking gate: if a leak is detected the recommendation is rejected, so the optimizer never hides a bug by enlarging a broken container. The pipeline combines SLA monitoring, a circuit breaker, HPA/PDB conflict detection, and a policy engine, with rollback and dry-run mode for human approval. Recommendations are produced by percentile analysis and Holt-Winters forecasting, balanced through multi-objective Pareto optimization at the per-container level. We validated the system with 1118 automated tests at 80.3% coverage and a live deployment on Google Kubernetes Engine, where right-sizing produced estimated cost savings of 20--40% in what-if projections and the leak gate reached 83% detection accuracy.

补充信息

↑