arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AdaptLSTM:分布漂移下云工作负载预测的高效自适应在线学习

AdaptLSTM: Efficient Adaptive Online Learning for Cloud Workload Forecasting under Distribution Drift

Xinhua Miao, Bowei Yang, Zhengong Cai

arXiv 2610.12265首次发表:更新:

发表机构

Zhejiang University(浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对云工作负载预测的分布漂移问题,提出自适应在线框架AdaptLSTM,在阿里云机器迹和容器迹数据集上以低计算成本实现了优异的预测性能,且与多种模型兼容。

AI 中文摘要

准确的工作负载预测对于网络规模云服务的弹性资源配置至关重要,而由病毒式内容、产品发布和用户行为驱动的分布偏移会快速降低离线训练模型的性能。朴素在线学习可恢复准确率,但会产生过高的每步计算成本。我们提出AdaptLSTM,一种自适应在线框架,通过验证校准阈值检测漂移并应用选择性、针对性更新。在阿里云机器迹数据集上,AdaptLSTM以20%的成本恢复了朴素在线学习54%的性能提升(效率提升2.7倍,10个随机种子的p值为0.002);在更具波动性的容器迹数据集上,它以20%的成本实现了96%的性能提升(效率提升4.8倍,较静态模型的平均绝对误差减少75%)。与经典漂移检测器ADWIN、DDM、Page-Hinkley(无法在回归规模误差流上触发)不同,AdaptLSTM在301步中触发42次,且优于匹配预算的基线。时钟分析显示其吞吐量提升1.33倍,更新时间减少45%。该框架与模型无关:在LSTM、GRU和Transformer主干上均呈现相同的帕累托模式。

英文摘要

Accurate workload forecasting is critical for elastic resource provisioning in web-scale cloud services, where distribution shifts driven by viral content, product launches, and user behavior degrade offline-trained models rapidly. Naive online learning recovers accuracy but incurs prohibitive per-step compute cost. We propose AdaptLSTM, an adaptive online framework that detects drift via validation-calibrated thresholds and applies selective, targeted updates. On the Alibaba Machine Trace, AdaptLSTM recovers 54\% of Naive Online's improvement at 20\% cost ($2.7\times$ efficiency, $p=0.002$ over 10 seeds). On the more volatile Container Trace, it achieves 96\% at 20\% cost ($4.8\times$ efficiency, $+75\%$ MAE reduction over Static). Unlike classical drift detectors (ADWIN, DDM, Page-Hinkley) which fail to trigger on regression-scale error streams, AdaptLSTM fires 42 times over 301 steps and outperforms matched-budget baselines. Wall-clock profiling shows $1.33\times$ throughput gain and 45\% update-time reduction. The framework is model-agnostic: identical Pareto patterns hold for LSTM, GRU, and Transformer backbones.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑