基于少样本连续上下文博弈的预测-校正循环用于需求预测
A Predict-then-Correct Loop Based on Few-Shot Continuous Contextual Bandit for Demand Forecasting
- Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
- Massachusetts Institute of Technology(麻省理工学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究针对零售需求预测难题,提出预测-校正框架,运用少样本连续上下文博弈校正策略等,经实验在多需求模式下显著降低误差、提高RMSE并降低库存成本,证明在线预测校正能连接离线需求学习与实时零售决策。
AI中文摘要:
当需求变化速度超过静态预测模型的重新训练速度时,零售需求预测仍然困难,尤其是在新观察标签稀疏的早期需求周期。为解决此问题,本研究提出预测-校正(PtC)框架,保留第一阶段机器学习预测,并应用少样本连续上下文博弈校正策略及相似SKU增强和top-p掩码更新。通过沃尔玛零售数据和独家饮料数据集实验,PtC在多种需求模式下显著降低MAPE、MAE和RMSE,消融研究中平均RMSE比仅用机器学习基线提高9.52%,且库存成本更低。结果表明在线预测校正可通过适应稀疏反馈连接离线需求学习和实时零售决策,而无需完全重新训练基础预测模型。
英文摘要:
Retail demand forecasting remains difficult when demand shifts faster than static forecasting models can be retrained, especially in early demand cycles where newly observed labels are sparse. To address this, this study aims to improve adaptive retail forecasting by proposing a predict-then-correct (PtC) framework that retains a first-stage machine learning (ML) forecast and applies a few-shot continuous contextual bandit correction policy with similar-SKUs augmentation and top-p masked updating. Across Walmart retail data and an exclusive beverage dataset, PtC delivers statistically significant reductions in MAPE, MAE, and RMSE across stable & high volume, stable & low volume, and erratic & intermittent demand patterns, improves average RMSE by 9.52% over the ML-only baseline in the ablation study, and yields lower inventory costs than base-stock, proximal policy optimization, and soft actor-critic policies under the tested lead-time settings. These findings show that online forecast correction can bridge offline demand learning and real-time retail decision-making by adapting to sparse feedback without fully retraining the base forecasting model.