发表机构
The University of Sydney(悉尼大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究采用归一化的问题级基准测试方法,对比6个开源和3个基线时间序列特征集在124个分类问题上的性能,发现整体表现相似,tsfresh性能最优,凸显特征组成与问题级性能的重要性。
AI 中文摘要
近年来,已开发出众多开源软件库用于从单变量时间序列中计算特征集。这些特征集的类型和数量各不相同,构建时采用了不同学科视角来量化时间序列数据的结构。迄今为止,这些特征集在时间序列分类问题上的相对优势和劣势仍未得到充分探索。本研究采用基于归一化的问题级基准测试方法(该方法相比以往基于排名的方法能更好地反映不同算法的相对优势与劣势),针对124个单变量时间序列分类问题,探究6个开源特征集和3个基线特征集(基于分布和/或基本频谱结构)的相对性能。尽管这些特征集在规模、组成和计算时间上存在显著差异,但整体表现较为相似:85.3%的成对比较结果为平局;其中规模最大的特征集tsfresh在所有与其他特征集的成对比较中,表现出最强的整体性能,胜率达29.03%。研究还指出了特定问题:在这些问题上,某一特征集的特定组成会使其具有显著的性能优势或劣势,而在另一些问题上,由傅里叶系数和分位数组成的简单基线足以实现良好性能。研究结果表明,在对时间序列特征集进行基准测试时,需考虑问题级性能,同时强调了特征组成对相对分类性能的重要驱动作用。
英文摘要
In recent years, numerous open-source software libraries have been developed for computing sets of features from univariate time series. The type and number of features vary across these feature sets, which have been constructed with varying disciplinary perspectives on quantifying structure in time-series data. To date, the relative strengths and weaknesses of these feature sets on time-series classification problems remains largely unexplored. Here we aimed to understand the relative performance of six open-source feature sets and three baseline feature sets (based on distributional and/or basic spectral structure) across 124 univariate time-series classification problems using a normalization-based approach to problem-level benchmarking that better indexes the relative strengths and weaknesses of different algorithms compared to prior rank-based approaches. Despite their dramatic differences in size, composition, and computation time, we found that feature sets performed relatively similarly overall (85.3% of pairwise comparisons resulted in ties), with the largest feature set, tsfresh, exhibiting the strongest overall performance (29.03% wins across all pairwise comparisons against other feature sets). We also highlighted specific problems on which the specific composition of a given feature set gave it a substantial performance advantage or disadvantage, and problems where simple baselines comprised of Fourier coefficients and quantiles were sufficient to achieve strong performance. Our results demonstrate the need to consider problem-level performance when benchmarking time-series feature sets, and highlight the importance of feature make-up in driving relative classification performance.
Comments24 pages, 3 figures