发表机构
The University of Melbourne; Banaras Hindu University(墨尔本大学; 贝拿勒斯印度教大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对云环境动态负载下反应式扩缩容响应延迟与振荡问题,提出基于双深度Q网络(DDQN+RRS)的主动扩缩容方法,通过解耦动作选择与价值评估提升稳定性,实验显示SLA违反率降至11.81%,CPU利用率达52.23%,扩缩容事件与Pod重启次数显著减少。
AI 中文摘要
动态工作负载和延迟敏感型应用要求在云计算环境中实现高效的自动扩缩容。然而,现有的大多数方法依赖于基于静态阈值的反应式机制,导致在工作负载不确定性下出现响应延迟和扩缩容振荡。为解决这些局限性,我们提出了一种基于双深度Q网络的主动式自动扩缩容方法(DDQN-Proactive),并辅以资源移除策略(RRS)。所提出的(DDQN+RRS)方法通过将动作选择与价值评估解耦来增强决策能力,从而实现更稳定和自适应的扩缩容。实验结果表明,所提出的方法优于反应式方法和现有的主动式方法。具体而言,DDQN+RRS实现了更低的服务水平协议(SLA)违反率(11.81%)、更高的CPU利用率(52.23%)、更优的扩缩容稳定性、更少的扩缩容事件(2,488次)以及更少的Pod重启次数(1,246次)。此外,该方法通过显著减少随时间(0-60秒)的振荡,确保了更平滑的自动扩缩容行为。虽然反应式方法在Pod分配上表现出显著的波动,但反应式方法减少了这些变化,而DDQN+RRS实现了最稳定和平滑的扩缩容,特别是在15-30秒、40-45秒和55-60秒的时间间隔内。
英文摘要
Dynamic workloads and latency-sensitive applications require efficient autoscaling in cloud computing environments. However, most existing approaches rely on reactive mechanisms based on static thresholds, resulting in delayed responses and scaling oscillations under workload uncertainty. To address these limitations, we propose a double deep Q-Network-based proactive autoscaling approach (DDQN-Proactive) along with Resource Removal Strategy (RRS). The proposed (DDQN+RRS) enhances decision-making by decoupling action selection from value evaluation, enabling more stable and adaptive scaling. Experimental results demonstrate that the proposed method outperforms both reactive and existing proactive approaches. Specifically, DDQN+RRS achieves a lower Service Level Agreement (SLA) violation rate (11.81%), higher CPU utilization (52.23%), improved scaling stability, fewer scaling events (2,488), and reduced pod restarts (1,246). Furthermore, the approach ensures smoother autoscaling behavior by significantly reducing oscillations over time (0-60 s). While reactive methods exhibit substantial fluctuations in pod allocation, Reactive reduces these variations, and DDQN+RRS achieves the most stable and smooth scaling, particularly during the 15-30 s, 40-45 s, and 55-60 s intervals.