arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

准确性并非服务:面向间歇性需求预测的决策感知基准

Accuracy Is Not Service: A Decision-Aware Benchmark for Intermittent-Demand Forecasting

Joo Ern Chin, Shih-Fen Cheng, Aldy Gunawan

arXiv 2609.13840首次发表:更新:

发表机构

Singapore Management University; ST Logistics Pte. Ltd.(新加坡管理大学; ST物流私人有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对间歇性需求预测,提出决策感知基准,发现预测准确性与订单服务负相关,偏差方向是关键,并给出无需训练的修正方法提升服务率。

AI 中文摘要

一家合同物流备件运营商按订单级服务获得报酬:只有当每个请求的行项都被满足时,订单才算完成,然而预测人员却是根据行项级预测准确性来选拔的。当需求具有间歇性和块状性、历史数据较短且提前期长达数月时,这种脱节就显得尤为重要。我们对38种预测方法进行了基准测试,涵盖经典方法、间歇性需求方法、机器学习、深度学习和预训练基础模型。一个通用的决策感知协议在来自实际合同和两个公开数据集的工业面板上对这些方法进行了评估。在工业面板上,针对20,330个真实多行项订单评估的方法中,预测准确性排名与订单服务排名呈负相关,相关系数为-0.555。服务与累积预测偏差的方向(包括零需求期间的过度预测)关联更紧密,而非与点准确性相关。检查Chronos-2实例归一化中的偏差,得出了一种无需训练的修正方法,在90%政策目标下,将每材料补货代理指标从77.5%提升至92.0%(提升14.5个百分点),并将完整订单补货率从54%提高到63%。为便于复现,我们发布了RUF(重新生成直至保真),一种生成保真认证合成面板的方法,在该面板上这些发现得以复现。对于间歇性需求,误差最低的预测未必能带来最高的服务。偏差方向有助于解释这一差距,且无需重新训练即可缩小该差距。

英文摘要

A contract-logistics spare-parts operator is paid on order-level service: an order counts only if every requested line is fulfilled, yet forecasters are selected based on line-level forecast accuracy. This disconnect matters when demand is intermittent and lumpy, histories are short, and lead times span months. We benchmarked 38 forecasting methods spanning classical, intermittent-demand, machine-learning, deep-learning, and pretrained foundation models. A common decision-aware protocol evaluates them on an industrial panel drawn from a live contract and two public datasets. Forecast-accuracy rank and order-service rank are negatively correlated on the industrial panel, at -0.555, across methods evaluated on 20,330 real multi-item orders. Service is more closely associated with the direction of cumulative forecast bias, including over-prediction during zero-demand periods, than with point accuracy. Examining bias in Chronos-2's instance normalization yields a training-free correction that lifts the per-material fill proxy from 77.5% to 92.0% (14.5 percentage points) at the 90% policy target and raises the complete-order fill rate from 54% to 63%. For reproducibility, we release RUF (Regenerate-Until-Fidelity), a method for generating fidelity-certified synthetic panels on which the findings reproduce. For intermittent demand, the lowest-error forecast need not deliver the highest service. Bias direction helps explain this gap, which can be reduced without retraining.

CommentsAccepted to the Twenty-Sixth IEEE International Conference on Data Mining (ICDM-26)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑