时间序列基础模型对工业监测是否有效?一项感知成本的实证研究
Do Time-Series Foundation Models Pay Off for Industrial Monitoring? A Cost-Aware Empirical Study
浏览论文内容
中文总结 AI 辅助
本研究通过三类工业监测场景对比评估时间序列基础模型与轻量基线的性能及成本,发现TSFMs并非拟合轻量模型的默认替代,而是任务依赖的部署选项。
中文摘要 AI 辅助
工业监测模型必须检测与操作相关的偏差,同时满足特定目标的数据、校准和资源约束。时间序列基础模型(TSFMs)有望提供可复用的表征和零样本预测,但在任务定义异质且轻量基线具有竞争力的情况下,其部署价值的证据仍参差不齐。本研究在三种场景下开展了感知协议的实证评估:以C-MAPSS作为退化风险代理、针对MIMII数据集的异常声音检测采用仅正常样本训练、以及带有合成目标扰动的BDG2预测残差诊断。我们从异常排序性能、风险- horizon敏感性、残差预测和扰动敏感性,以及本地实现成本等方面,评估了经典单类方法、紧凑神经自编码器、残差预测器、MOMENT-small、Chronos-T5和TimesFM 2.5。在交叉验证的100台C-MAPSS发动机中,TCN-AE的交叉验证加权AUROC/AUPRC达到0.9570/0.8960,而MOMENT重构的对应值为0.7310/0.3080;成对发动机聚类自举置信区间显示两者差异不为零。在5个匹配的MIMII泵评估中,OCSVM的AUROC和AUPRC也超过MOMENT重构。在固定12米BDG2面板上,TimesFM 2.5的对齐预测误差最低,合成AUROC点估计最高,不过TSFM与拟合残差模型的合成AUPRC相近。同设备测量显示,MOMENT的延迟、峰值分配VRAM和序列化状态字典大小均高于TCN-AE。在评估的冻结和零样本设置下,TSFMs是依赖任务的部署选项,而非拟合轻量模型的默认替代品。
英文摘要
Industrial monitoring models must detect operationally relevant deviations while satisfying target-specific data, calibration, and resource constraints. Time-series foundation models (TSFMs) promise reusable representations and zero-shot forecasts, yet evidence for their deployment value remains mixed when task definitions are heterogeneous and lightweight baselines are competitive. This work presents a protocol-aware empirical assessment across three settings: a C-MAPSS degradation-risk proxy, normal-only training for anomalous-sound detection on MIMII, and BDG2 forecasting-residual diagnostics with synthetic target perturbations. We assess classical one-class methods, compact neural autoencoders, residual forecasters, MOMENT-small, Chronos-T5, and TimesFM 2.5 in terms of anomaly-ranking performance, risk-horizon sensitivity, residual forecasting and perturbation sensitivity, and local implementation cost. Across 100 C-MAPSS engines evaluated out of fold, TCN-AE reaches fold-weighted AUROC/AUPRC 0.9570/0.8960, compared with 0.7310/0.3080 for MOMENT reconstruction; paired engine-cluster bootstrap confidence intervals exclude zero for both differences. Across five matched MIMII pump evaluations, OCSVM also exceeds MOMENT reconstruction in AUROC and AUPRC. On a fixed 12-meter BDG2 panel, TimesFM 2.5 has the lowest aligned forecast error and the highest synthetic AUROC point estimate, although synthetic AUPRC is similar across TSFM and fitted residual models. Same-device measurements show that MOMENT incurs higher latency, peak allocated VRAM, and serialized state-dictionary size than TCN-AE. Under the evaluated frozen and zero-shot settings, TSFMs are task-dependent deployment options rather than default replacements for fitted lightweight models.
发表机构
- National Taiwan University of Science and Technology(台湾科技大学)
机构由 AI 辅助整理,请以论文原文为准。