发表机构
Univ Rennes; Ensai; CNRS; CREST(雷恩大学; 国立统计与数据分析学校; 法国国家科学研究中心; 经济与统计研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
TIGER通过将三个通用分类器应用于四种表示族,堆叠预测并用自适应规则选择加权投票或表格基础模型,在142个数据集上取得最优性能。
AI 中文摘要
表示族是从时间序列中提取特征的不同方式。集成多种表示族的集成算法仍然是时间序列分类最准确的方法。当前最先进的集成方法,尤其是HIVE-COTE 2.0,为每个表示族配对专门的分类算法,并使用固定的、非自适应的规则组合它们的预测。我们提出TIGER(基于上下文学习门控表示集成的时序分类),它应用相同的小型通用分类器组合(Ridge、Extra Trees和朴素贝叶斯)到来自四个不同表示族的四种表示上,将得到的十二个基学习器的预测堆叠成元特征矩阵。最终预测由自适应元分类规则产生,该规则针对每个数据集独立地在加权硬多数投票和TabICLv2(一种预训练的表格基础模型,用作上下文元分类器)之间进行选择,基于每类可用训练样本的平均数量。在从UCR时间序列分类档案中抽取的142个数据集的基准上,TIGER在六种比较算法(包括HIVE-COTE 2.0)中获得了最佳的平均准确率、平衡准确率和F1分数,并显著优于其他五种算法中的每一种。TIGER的自适应规则也显著优于其两种组成元分类方法单独使用时的表现,其单一超参数仅使用二十个数据集的开发子集进行调优,并显示出对完整评估基准的泛化能力。我们进一步通过大量消融实验来表征TIGER的设计,并报告了我们研究并最终放弃的设计替代方案。
英文摘要
A representation family is a distinct way of extracting features from time series. Ensemble algorithms that combine several representation families remain the most accurate approach to time series classification. Current state-of-the-art ensembles, most notably HIVE-COTE~2.0, pair a bespoke classification algorithm with each representation family and combine their predictions using a fixed, non-adaptive rule. We present TIGER (Time-series classification with In-context-learning Gated Ensemble of Representations), which instead applies the same small portfolio of three general-purpose classifiers (Ridge, Extra Trees, and Naive Bayes) to four representations from four distinct families, stacking the resulting twelve base learners' predictions into a meta-feature matrix. The final prediction is produced by an adaptive meta-classification rule that chooses, independently for each data set, between a weighted hard majority vote and TabICLv2, a pretrained tabular foundation model used in-context as a meta-classifier, based on the mean number of training samples available per class. On a 142-data-set benchmark drawn from the UCR time series classification archive, TIGER obtains the best mean accuracy, balanced accuracy, and F1-score among six compared algorithms, including HIVE-COTE~2.0, and significantly outperforms each of the other five individually. TIGER's adaptive rule also meaningfully outperforms either of its two constituent meta-classification methods used alone, and its single hyperparameter, tuned using only a twenty-data-set development subset, is shown to generalize to the full evaluation benchmark. We further characterize TIGER's design through an extensive set of ablation experiments and report the design alternatives that we investigated and ultimately discarded.