arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.40117cs.LG

超越模型排名:工业时间序列预测中分布-统计错误设定的机制诊断

Beyond Model Ranking: Regime Diagnosis for Distributional-Statistical Misspecification in Industrial Time-Series Forecasting

Pengyu Nie, Chenglang Xu, Yaoshi Chen, Chaogan Ren, Wei Hu, Chao Yang, Jiangong Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

针对工业时间序列预测中训练与部署偏差,提出机制相对偏差向量(RBV)诊断工具,区分损失与模型导致的偏差,并证明机制感知训练可解决池化诱导偏差,补充模型排名。

中文摘要 AI 辅助

时间序列预测模型在基准测试中表现优异,但在工业部署中却表现出严重的系统性偏差。这种训练与部署之间的差距通常被归因于时间结构错误或分布偏移。我们刻画了一个这些解释所忽视的互补性来源:标准损失函数嵌入了固定的统计先验,而工业需求混合了良性及病态机制——零膨胀、偏态、高变异性——在这些机制中,这些先验被系统性违反。即使在完美的时间建模下,所引发的偏差依然存在,并保留在归一化无法去除的分布形状成分中,同时产生聚合指标无法显现的聚合权衡。我们将这些观察转化为一个以机制相对偏差向量(RBV)为核心的评估工具包:一种与度量无关、按机制分解的诊断方法,用于审计池化训练如何在病态子群体间分配系统性错配。一项受控归因分析将RBV分解为与模型无关的内在下限(由每种损失的估计目标设定)和可归因于训练的额外成分,从而将观察到的偏差追溯到损失而非模型。一项大规模研究——涵盖13种损失目标、3个种子、跨越RetailShiftBench和M5的60,000多个序列,并带有随机分割对照——表明,机制感知诊断能够区分优化型失败与偏差型失败,并且对于均值型损失,机制感知训练能够解决容量扩展无法解决的池化诱导偏差。一项正式的结构性观察——即评估分布污染下的风险是病态混合权重的仿射函数——为这些发现提供了基础。我们的工作以机制导向、基于机制的评价补充了模型排名。

英文摘要

Time-series forecasting models achieve strong benchmark performance but exhibit severe systematic bias in industrial deployments. This train--deploy gap is conventionally attributed to temporal-structural errors or distribution shifts. We characterize a complementary source that these explanations overlook: canonical losses embed fixed statistical priors, while industrial demand mixes benign and pathological regimes---zero-inflation, skewness, high variability---in which these priors are systematically violated. The induced bias persists even under perfect temporal modeling, remains in a distributional-shape component that normalization cannot remove, and creates an aggregation trade-off invisible to aggregate metrics. We turn these observations into an evaluation toolkit centered on the Regime-wise Relative Bias Vector (RBV): a metric-agnostic, regime-decomposed diagnostic that audits how pooled training allocates systematic mismatch across pathological subpopulations. A controlled attribution analysis decomposes RBV into a model-independent intrinsic floor, set by each loss's estimand, and an excess component attributable to training, tracing observed bias to the loss rather than the model. A large-scale study---13 loss objectives, 3 seeds, 60,000+ series spanning RetailShiftBench and M5, with random-split controls---shows that regime-aware diagnosis separates optimization-type from bias-type failure, and that regime-aware training resolves the pooling-induced bias that capacity scaling cannot, for mean-type losses. A formal structural observation, that risk under evaluation-distribution contamination is affine in the pathology mixture weight, grounds these findings. Our work complements model ranking with mechanism-grounded, regime-oriented evaluation.

发表机构

  • JD.com, Inc.(京东集团)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑