Argus:一项关于预测Spot中断何时优于简单检查点的真实EKS研究
Argus: A Real-EKS Study of When Predicting Spot Interruptions Beats Simple Checkpointing
浏览论文内容
中文总结 AI 辅助
本研究构建Kubernetes操作器Argus,在真实EKS和CIFAR-10上实证比较预测中断与简单检查点,发现小提前量预测可零浪费,但过大提前量导致过度迁移,为Spot中断应对提供指导。
中文摘要 AI 辅助
弹性计算云(EC2)Spot实例比按需实例便宜60%至90%,但可能仅提前2分钟通知即被回收;对于昂贵的多节点训练,这种损失可能很严重,一次回收可能耗费数小时的同步训练进度。我们构建了Argus,一个Kubernetes操作器,并在CIFAR-10测试平台上实证研究预测中断何时优于简单检查点。Argus在真实的EKS上通过优雅的SIGTERM检查点成功应对真实的Spot排空,从第8轮恢复,仅丢失正在进行的轮次。此外,我们在80次试验的基准测试中发现,当中断速度超过固定的2分钟通知时间时,基于通知的响应策略会退化至无保护状态,而预测浪费的计算量降至零;但若固定提前量过大,则过度迁移严重,在最快速度下五次运行仅有一次完成,而周期性检查点是一个无需机器学习的强基线。提前量扫描将提前预测转化为指导原则:较小的提前量足以实现零浪费,但过大的提前量则造成浪费。所构建的预测器是建议性的(代理标签);真实中断标签和大规模模型验证是未来工作。
英文摘要
Elastic Compute Cloud (EC2) Spot is 60% to 90% cheaper than On-Demand but can be reclaimed on just a 2-minute notice; for expensive multi-node training this loss can be severe, with one reclaim costing hours of synchronous progress. We build Argus, a Kubernetes operator, and ask empirically, on a CIFAR-10 testbed, when predicting interruptions beats simple checkpointing. Argus on real EKS survives a real Spot drain with a graceful SIGTERM checkpoint, resuming from epoch 8 and losing only the in-progress epoch. Alongside, we further find that in an 80-trial benchmark, the reactive-on-notice degrades toward no protection once interruption outpaces the fixed 2-minute notice, and predictive wasted compute is driven to zero, but with an oversized fixed lead it over-migrates so severely that at the fastest rate only one of five runs completes, while periodic is a strong ML-free baseline. A lead-time sweep turns the lead prediction into a guideline where a small lead suffices for zero waste, but excess lead is wasteful. The predictor built is advisory (a proxy label); real interruption labels and large-model-scale validation are future work.