arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.19936cs.DC

VERA:Kubernetes中HPC工作负载动态内存扩展的强化学习

VERA: Reinforcement Learning for Dynamic Memory Scaling of HPC Workloads in Kubernetes

Ade Pramono, Jie Ren, Ivy Peng

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出基于强化学习的VERA推荐器,将Kubernetes中HPC工作负载的垂直内存扩展建模为马尔可夫决策过程,在真实轨迹上训练,实验表明其可回收31.6%内存余量并减少OOM事件,优于默认VPA。

中文摘要 AI 辅助

当HPC工作负载在Kubernetes上运行时,内存过度配置会导致资源利用率不足。默认的Vertical Pod Autoscaler(VPA)无法预测首次运行的HPC作业的阶段性内存峰值。在这项工作中,我们提出了一种强化学习(RL)推荐器VERA,它将垂直内存扩展形式化为马尔可夫决策过程,并在3353条真实Prometheus轨迹上训练了一个智能体。在实时Google Kubernetes Engine集群上使用LAMMPS、图分析、内存分析和MLPerf 3D-UNet进行评估,该RL智能体回收了可用内存余量的31.6%,并且最多发生一次OOM事件,而VPA在相同运行中回收了-7.9%,提高了内存配置,其建议在30次运行中不足以避免OOM。结果表明,基于观察的RL推荐器在动态内存扩展方面可以优于回顾性启发式方法。

英文摘要

Memory over-provisioning results in resource underutilization when HPC workloads run on Kubernetes. The default Vertical Pod Autoscaler (VPA) cannot anticipate phase-driven memory spikes for first-run HPC jobs. In this work, we present a reinforcement learning (RL) recommender VERA that formulates vertical memory scaling as a Markov Decision Process and trains an agent on 3353 real Prometheus traces. Evaluated on a live Google Kubernetes Engine cluster using LAMMPS, graph analytics, in-memory analytics, and MLPerf 3D-UNet, the RL agent reclaims 31.6% of the available memory headroom and incurs at most one OOM event while VPA reclaims -7.9% over the same runs, raising memory provisioning, and its recommendation would have been insufficient to avoid OOM in 30 runs. The results demonstrate that an observation-driven RL recommender could outperform retrospective heuristics for dynamic memory scaling.

发表机构

  • KTH Royal Institute of Technology(皇家理工学院)
  • William & Mary(威廉与玛丽学院)

机构由 AI 辅助整理,请以论文原文为准。

↑