arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.15148cs.LG

基于Shapelet距离度量的时间序列多源迁移学习

Multi-source Transfer Learning of Time Series with a Shapelet-based Distance Measure

  • Kyushu University(九州大学)

机构由 AI 辅助整理,请以论文原文为准。

Jiseok Lee, Brian Kenji Iwana

AI总结:

本文提出一种无需训练的Shapelet Matching源选择方法,通过shapelet相似度选择多个源数据集进行多源预训练,在UCR 128个数据集上降低了负迁移风险,并在CNN和Transformer架构上表现优异。

AI中文摘要:

迁移学习是解决时间序列分类深度学习中的数据稀缺问题的有效技术,但其成功取决于源数据集的选择。传统的可迁移性估计方法通常计算成本高昂,因为它们需要对每个潜在的源数据集进行完整的模型预训练以评估其适用性。本文提出了一种新颖的、无需训练(免训练)的源选择方法,名为Shapelet Matching(形状匹配)。我们的方法首先从目标数据集和潜在的源数据集中识别出具有判别性的shapelets(形状特征)。然后,Shapelet Matching通过比较提取的shapelet集合来量化数据集之间的相似性。为了减轻因选择不合适的单一源而导致的负迁移风险,我们引入了一种多源迁移学习方法。我们根据基于shapelet的相似性得分选择多个源数据集,将它们合并为一个单一的多源数据集,并使用该聚合数据集进行预训练。随后,模型在目标任务上进行微调。我们使用时间卷积神经网络(CNN)和Transformer架构,在UCR Archive的128个数据集上评估了我们的方法。实证结果表明,我们的多源预训练平均降低了负迁移的风险。Shapelet Matching在CNN骨干网络上取得了最强的性能,并且在基于补丁的Transformer架构上保持竞争力,同时避免了为每个候选源单独预训练模型的成本。

英文摘要:

Transfer learning is an effective technique for addressing data scarcity in deep learning for time series classification, but its success depends on the selection of source datasets. Conventional transferability estimation methods are often computationally expensive, as they require fully pre-training a model on each potential source dataset to assess its suitability. This paper introduces a novel, training-free source selection method named Shapelet Matching. Our approach first identifies discriminative shapelets from the target and potential source datasets. Then, Shapelet Matching quantifies dataset similarity by comparing the extracted sets of shapelets. To mitigate the risk of negative transfer from selecting an unsuitable single source, we introduce a multi-source transfer learning method. We select several source datasets based on their shapelet-based similarity scores, combine them into a single multi-source dataset, and use this aggregated dataset for pre-training. The model is then fine-tuned on the target task. We evaluated our method on 128 datasets from the UCR Archive using both temporal CNN and Transformer architectures. The empirical results demonstrate that our multi-source pre-training reduces the risk of negative transfer on average. Shapelet Matching achieves the strongest performance for the CNN backbone and remains competitive for patch-based Transformer architectures, while avoiding the cost of pre-training a separate model for every candidate source.

补充信息

↑