再探WEASEL 2.0:复现、敏感性分析与自适应集成规模规则
Revisiting WEASEL 2.0: Reproduction, Sensitivity, and an Adaptive Ensemble-Size Rule
浏览论文内容
中文总结 AI 辅助
本研究复现了WEASEL 2.0,分析其四项设计选择的敏感性,发现最大集成规模规则对长序列数据集配置过度,提出自适应规则后可减少内存和时间消耗且准确率影响极小。
中文摘要 AI 辅助
WEASEL 2.0是一种基于字典的时间序列分类器,它结合了扩张滑动窗口、随机超参数集成以及固定大小的密集特征表示。其两个超参数选择项——最大集成规模和最大窗口大小,由简单阈值规则指定,而所选阈值在原始论文中未得到经验验证。本研究在114个UCR数据集上复现了WEASEL 2.0,获得了0.865的平均准确率和0.928的中位数准确率,与已发表值高度匹配(Wilcoxon符号秩检验,p=0.655)。随后,测试了四项设计选择的敏感性:下游分类器、特征权重缺失、最大窗口大小规则以及最大集成规模规则。前三项对扰动具有鲁棒性,第四项对于长序列数据集而言配置过度,这促使我们提出一种自适应规则,根据序列长度和类别数来设置最大集成规模。在定长数据集上评估时,该自适应规则使峰值拟合内存减少了37 MB(中位数,均值为395 MB),拟合时间减少了0.4秒(中位数,均值为4秒),准确率变化的中位数为0%(均值为-0.11%)。内存和时间的节省集中在长序列数据集上,而原始规则在这些数据集上分配了最大的集成规模。
英文摘要
WEASEL 2.0 is a dictionary-based time series classifier that combines dilated sliding windows with a randomised hyperparameter ensemble and a fixed-size dense feature representation. Two of its hyperparameter choices, the maximum ensemble size and the maximum window size, are specified by simple thresholding rules whose chosen thresholds are not empirically justified in the original paper. In this work we reproduce WEASEL 2.0 on 114 UCR datasets, achieving a mean accuracy of 0.865 and median of 0.928, closely matching the published values (Wilcoxon signed-rank, p = 0.655). We then test the sensitivity of four design choices: the downstream classifier, the absence of feature weighting, the maximum window-size rule, and the maximum ensemble-size rule. The first three are robust to perturbation. The fourth is over-provisioned for long-series datasets, motivating an adaptive rule that sets the maximum ensemble size from series length and number of classes. Evaluated on fixed-length datasets, the adaptive rule reduces peak fit memory by a median of 37 MB (mean 395 MB) and fit time by a median of 0.4 s (mean 4 s), with a median accuracy change of 0% (mean -0.11%). Memory and time savings concentrate on long-series datasets where the original rule allocates the largest ensemble size.
发表机构
- University College Dublin(都柏林大学学院)
机构由 AI 辅助整理,请以论文原文为准。