发表机构
University of Texas at Dallas; Neuralix AI; University of California, Riverside(德克萨斯大学达拉斯分校; Neuralix AI; 加州大学河滨分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出名为MINT的张量分解方法,将其应用于堆叠递归矩阵,在多类时间序列数据中有效实现跨传感器模式的共聚类。
AI 中文摘要
递归图是一种应用于多个领域(如恒星光变曲线、声音波形、CCT遥测)的时间序列数据挖掘基础工具。本研究提出张量化自相似矩阵作为N条长度为n的单变量时间序列数据集(N×n)的基础工具,其子序列窗口长度为m,且其基于张量的特性可自然扩展至多变量数据集。所提计算该基础工具的方法会从这些数据集中计算出大小为N×(n−m+1)×(n−m+1)的点积图,随后使用张量分解方法对该张量进行挖掘以发现共聚类模式。我们在大规模快速交通、电力需求、风力涡轮机及汽车流量数据中验证了结果,发现MINT流程可在包含规则间隔模式的高度规则数据中有效对跨传感器模式进行共聚类。
英文摘要
Recurrence plots are a time series data mining primitive applied to a variety of domains (e.g. star light curves, sound waveforms, CCT telemetry). This work proposes tensorized self-similarity matrices as a primitive for univariate time series datasets ($N\times n$) of $N$ time series of length $n$ with a subsequence window of length $m$, and whose tensor-based nature is naturally extensible to multivariate datasets. The proposed method to compute this primitive computes dot plots of size $N \times (n-m+1) \times (n-m+ 1)$ from these datasets, where the subsequent tensor is mined using tensor decomposition methods to mine for co-clustered patterns. We demonstrate our results in mass rapid transit, electricity demand, wind turbine, and car traffic data, finding the MINT pipeline effectively co-clusters cross-sensor patterns in highly regular datasets containing motifs at regular intervals.