arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27867cs.LGstat.APstat.ME

评估选择决定预测排行榜:来自生产市场面板的证据

Evaluation Choices Decide the Forecasting Leaderboard: Evidence from a Production Marketplace Panel

Md Rezwanul Islam, Wael Mohammed

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过在生产市场面板上对24种预测方法进行基准测试,证明评估设计选择(如分析单位、误差汇总和评分方式)能逆转或消解预测排行榜结论,并量化了部署系统选择规则的效果。

中文摘要 AI 辅助

预测基准报告了哪种方法获胜。我们表明,答案在任何模型拟合之前就由评估者的选择决定。我们在一个包含1,887个商业客户、跨越67个月的生产市场面板上,对24种预测方法和一个教科书参考方法(包括六个2025时代的时间序列基础模型)进行了基准测试。我们固定数据、预测期限和时期,仅改变评估设计。三个选择各自逆转或消解了一个头条结论。将分析单位从市场总量改为单个客户,使我们的生产基线从十九种方法中的第二名(未被任何方法击败)变为二十五种方法中的第二十三名。在后者中,其二十四个挑战者中有十九个击败了它。改变误差汇总的程度决定了Diebold-Mariano检验是否能发现任何显著差异。对预测区间而非点预测进行评分几乎完全重新排序了领域,在间歇性需求上的秩相关系数为0.02。然后我们衡量了部署系统从中获得的效果。其选择规则捕获了从不采取行动到事后选择之间距离的55%。这种逆转并非我们数据的特例。我们在公开的M5零售面板上原封不动地运行了已发布的协议。相同的基线形态在市场总量上排名第一,而在每个序列上排名最后,被所有方法击败,重放的选择规则在那里关闭了相同从地板到天花板距离的64.7%。在该名单中添加五个零样本基础模型改变了在总量上的获胜者,但没有改变形态。预测区间的盲点也随行:共形区间在峰值最高的项目上覆盖不足。将我们自己的面板分割成更小的组,将对比变成了一条曲线:基线的排名在每一级分解中都会恶化。我们发布了评估协议,并报告了我们自己的一个错误,该错误在我们发现之前逆转了一个结果。

英文摘要

A forecasting benchmark reports which method won. We show that the answer is set by the evaluator's choices before any model is fitted. We benchmark 24 forecasting methods and one textbook reference, including six 2025-era time series foundation models, on a production marketplace panel of 1,887 business customers over 67 months. We hold the data, the horizon and the period fixed, and vary only the evaluation design. Three choices each reverse or dissolve a headline conclusion. Changing the unit of analysis from the market total to the individual customer moves our production baseline from second of nineteen, beaten by nothing, to twenty-third of twenty-five. Nineteen of its twenty-four challengers beat it there. Changing how much error is pooled decides whether a Diebold-Mariano test finds anything at all. Scoring prediction intervals rather than point forecasts reorders the field almost completely, with a rank correlation of 0.02 on intermittent demand. We then measure what the deployed system gets from this. Its selection rule captures 55% of the distance between doing nothing and choosing with hindsight. The reversal is not a quirk of our data. We ran the released protocol, unchanged, on the public M5 retail panel. The same baseline shape places first at the market total and last per series, beaten by everything, and a replayed selection rule closes 64.7% of the same floor-to-ceiling distance there. Adding five zero-shot foundation models to that roster changes who wins at the total, not the shape. The bands' blind spot travels too: conformal bands under-cover most on the spikiest items. Splitting our own panel into ever smaller groups turns the contrast into a curve: the baseline's rank worsens at every level of disaggregation. We release the evaluation protocol and report an error of our own that inverted a result before we caught it.

发表机构

  • Field Nation LLC(Field Nation有限责任公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑