发表机构
Amazon(亚马逊)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过两个真实库存控制部署,验证了反事实模拟器推演训练预测模型可解决新策略冷启动问题,显著降低MAPE误差,且后续校准可进一步优化。
AI 中文摘要
部署新的决策策略会给预测模型带来冷启动问题,这些模型的目标依赖于策略的行动:历史观测反映的是早期策略,而新策略下的真实观测尚不可用。模拟提供了一种弥补这一差距的方法,即在反事实场景中推演目标策略,并利用生成的轨迹来学习系统对这些控制的响应。这种模拟器训练模型的模拟到现实(Sim2Real)迁移,可以通过将其与过去部署中的真实观测进行对比评估来进行回测。利用两个真实世界的库存控制部署,我们从三个角度评估了这一过程:模拟器保真度、对真实行为的零样本迁移,以及随着真实目标策略观测的积累而进行的适应。模拟器训练的预测器在点估计平均绝对百分比误差(MAPE)上低于同一架构在历史真实数据上训练的模型,在研究1中MAPE降低了1.2-3.1个百分点,在研究2中降低了12.5-18.7个百分点。部署后,利用早期真实观测进行轻量级校准进一步将误差降低了最多2.5个百分点。这些结果提供了实证证据,表明模拟器生成的反事实数据可以支持新策略下的冷启动预测,并且随着真实部署数据的可用,所得模型可以进一步优化。
英文摘要
Deploying a new decision policy creates a cold-start problem for prediction models whose targets depend on the policy's actions: historical observations reflect earlier policies, while real observations under the new policy are not yet available. Simulation offers a way to address this gap by rolling out the target policy across counterfactual scenarios and using the resulting trajectories to learn how the system responds to those controls. The simulation-to-reality (Sim2Real) transfer of this simulator-trained model can then be backtested by evaluating it against real observations from past deployments. Using two real-world inventory-control deployments, we evaluate this process from three angles: simulator fidelity, zero-shot transfer to real behavior, and adaptation as real target-policy observations accumulate. The simulator-trained forecaster achieves lower point-estimate mean absolute percentage error (MAPE) than the same architecture trained on historical real data, reducing MAPE by 1.2-3.1 percentage points in Study 1 and 12.5-18.7 points in Study 2. After deployment, lightweight calibration using early real observations further reduces error by up to 2.5 percentage points. These results provide empirical evidence that simulator-generated counterfactual data can support cold-start forecasting under a new policy, and the resulting model can be further refined as real deployment data become available.
Comments15 pages, 3 figures, 9 tables