arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

机器学习何时能超越价值排序?对暴露加权发货优先级的三个数据集诊断

When Does Machine Learning Beat Value Sorting? A Three-Dataset Diagnostic of Exposure-Weighted Shipment Prioritization

Jize Li

arXiv 2607.18573首次发表:更新:

AI 中文总结

研究在特定供应链场景下机器学习能否超越价值排序,通过泄漏控制的滚动原点评估等方法,发现按预测延迟严重性乘以已知价值排序在部分数据集优于仅按严重性排序,但未普遍超越价值排序,提出部署诊断和评估协议。

AI 中文摘要

延迟风险模型通常通过预测准确性来判断。但在实际中,重要的是在只能审查少量发货的情况下,经理应首先检查哪些发货。我们评估机器学习是否能超越一个严格的无模型基线:首先检查最高价值的发货。在三个真实供应链场景(SCMS采购、DataCo物流和Olist电子商务)中,我们使用泄漏控制的滚动原点评估和1000样本配对自举置信区间。按预测延迟严重性乘以已知价值排序(M1)在所有三个数据集中都优于仅按严重性排序,但通常无法超越价值排序。在10%的审查预算下,SCMS中M1减去仅按价值排序为-5.5个百分点,DataCo为+10.1个百分点,Olist为-4.9个百分点。这种差异与严重性可学习性一致:DataCo的R^2 = 0.但在实际中,重要的是在只能审查少量发货的情况下,经理应首先检查哪些发货。我们评估机器学习是否能超越一个严格的无模型基线:首先检查最高价值的发货。在三个真实供应链场景(SCMS采购、DataCo物流和Olist电子商务)中,我们使用泄漏控制的滚动原点评估和1000样本配对自举置信区间。按预测延迟严重性乘以已知价值排序(M1)在所有三个数据集中都优于仅按严重性排序,但通常无法超越价值排序。在10%的审查预算下,SCMS中M1减去仅按价值排序为-5.5个百分点,DataCo为+10.1个百分点,Olist为-4.9个百分点。这种差异与严重性可学习性一致:DataCo的R^2 = 0.27,校准偏差为+0.01天,而SCMS和Olist的R^2约为-0.02,校准偏差为负。嵌套交叉验证成本敏感再训练并未比M1带来稳定改进。本文提出了一个部署诊断和评估协议。价值排序应始终是一个基准,只有在对严重性可学习性和校准进行审核且模型在泄漏控制的滚动原点评估下通过该关卡后,才应部署机器学习。

英文摘要

Delay-risk models are usually judged by predictive accuracy. What matters in practice is narrower: with capacity to review only a few shipments, which ones should a manager check first? We evaluate whether machine learning clears a demanding no-model baseline: inspect the highest-value shipments first. Across three real supply-chain contexts: SCMS procurement, DataCo logistics, and Olist e-commerce, we use leakage-controlled rolling-origin evaluation and 1000-sample paired bootstrap confidence intervals. Ranking by predicted delay severity times known value (M1) beats severity-only ranking in all three datasets, yet it does not generally beat value sorting. At a 10% review budget, M1 minus VALUE_ONLY is -5.5 percentage points (pp) for SCMS, +10.1 pp for DataCo, and -4.9 pp for Olist. The divide is consistent with severity learnability: DataCo has R^2 = 0.27 and calibration bias of +0.01 days, whereas SCMS and Olist have R^2 of approximately -0.02 and negative calibration bias. Nested-CV cost-sensitive retraining does not deliver a stable improvement over M1. Rather than proposing a new learning algorithm, this paper presents a deployment diagnostic and evaluation protocol. Value sorting should remain a permanent benchmark, and ML should be deployed only after severity learnability and calibration have been audited and the model clears that gate under leakage-controlled rolling-origin evaluation.

Comments9 pages, 1 figure, 4 tables, accepted for presentation at AIIIP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑