arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

缺失模式不匹配对时间序列分类方法选择的影响:一项受控实证研究

The Effect of Missingness-Pattern Mismatch on Method Selection for Time-Series Classification: A Controlled Empirical Study

Ruiqi Zhao, Zishun Yuan, Zhentao Wang, Jiahao Quan, Kangzheng Li, Jianfan Deng

arXiv 2610.05368首次发表:更新:

发表机构

The University of Tokyo; Stony Brook University, SUNY Korea; Fuyao University of Science and Technology; Renmin University of China; Tsinghua University; Anhui Science and Technology University(东京大学; 纽约州立大学石溪分校韩国校区; 福耀科技大学; 中国人民大学; 清华大学; 安徽科技学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过受控实验发现,验证与部署数据缺失模式不匹配会降低所选时间序列分类器的平均测试准确率,匹配缺失模式是方法选择的简单保障。

AI 中文摘要

时间序列分类的分类器通常在验证数据上进行选择,但部署时缺失观测的时间模式可能与验证期间看到的模式不同。我们研究了这种不匹配是否会影响基于验证的分类器选择。在一项受控的 $2 \ imes 2$ 设计中,对64个单变量UCR数据集的验证集和测试集以5%至30%的六个比率,分别使用随机点缺失或循环块缺失进行掩蔽,通过线性插值进行填补,并用于在三个预选候选方法中进行选择:1NN-DTW、带有Ridge分类器的MiniRocket以及基于统计特征的随机森林。训练数据保持完整,并在相同的掩蔽测试集上比较了在匹配和不匹配验证模式下所做的选择。不匹配的验证使所选分类器的测试平衡准确率平均降低了1.14个百分点(95%置信区间0.79至1.51),在64个数据集中有49个出现损失。在5%缺失率时损失可忽略不计,在30%时增加到2.46个百分点。损失集中在点掩蔽部署(1.84个百分点),其中块掩蔽验证使选择偏离了通常最佳的候选方法,而块掩蔽部署的影响较小且不显著。不匹配在35.5%的配对比较中改变了所选分类器,但改变选择并不总是降低性能。一项使用非包裹线性块的补充分析重现了这些发现,且效应更大(1.67个百分点)。因此,将验证数据的缺失模式与预期的部署模式相匹配是方法选择的一个简单保障,尤其是在较高的缺失率下。

英文摘要

Classifiers for time-series classification are commonly selected on validation data, but the temporal pattern of missing observations at deployment may differ from the pattern seen during validation. We examine whether such a mismatch affects validation-based classifier selection. In a controlled $2 \times 2$ design, validation and test sets of 64 univariate UCR datasets were masked with either random point missingness or circular block missingness at six rates from 5% to 30%, imputed by linear interpolation, and used to select among three prespecified candidates: 1NN-DTW, MiniRocket with a Ridge classifier, and a statistical-feature Random Forest. Training data remained complete, and selections made under matched and mismatched validation patterns were compared on the same masked test sets. Mismatched validation reduced the test balanced accuracy of the selected classifier by 1.14 percentage points on average (95% CI 0.79 to 1.51), with losses on 49 of the 64 datasets. The loss was negligible at 5% missingness and increased to 2.46 percentage points at 30%. It was concentrated in point-masked deployment (1.84 percentage points), where block-masked validation shifted selection away from the usually best candidate, while the effect for block-masked deployment was small and not significant. Mismatch changed the selected classifier in 35.5% of paired comparisons, but a changed selection did not always reduce performance. A supplementary analysis with non-wrapping linear blocks reproduced these findings with a larger effect (1.67 percentage points). Matching the missingness pattern of validation data to the expected deployment pattern is therefore a simple safeguard for method selection, particularly at higher missingness rates.

Comments26 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑