arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Panache:在每个窗口长度下进行单遍基序发现

Panache: One-Pass Motif Discovery at Every Window Length

Tej Sanibh Ranade

arXiv 2607.17481首次发表:更新:

AI 中文总结

研究针对时间序列基序发现中模式持续时间未知的问题,提出单遍流算法Panache。通过滑动DFT递归等维护频谱状态,利用Parseval定理在哈希目录中让相似子序列碰撞,自行计算依赖数据的参数,在多个数据集上表现出色,速度远超基线方法。

AI 中文摘要

基序发现是探索性数据分析的核心要素,旨在在时间序列中寻找重复模式。但模式的持续时间通常事先未知,现有方法需在定义的窗口长度区间内逐一尝试。现有的泛矩阵轮廓(PMP)方法每计算一个长度的z归一化矩阵轮廓,就要对同一序列进行一次二次自连接,计算成本高。本文引入Panache,这是首个用于z归一化PMP基序发现的单遍流算法,其运行时间与序列长度近似线性关系。通过滑动离散傅里叶变换递归和运行统计在线维护每个z归一化子序列的非直流频谱,利用Parseval定理在占用控制哈希目录中让相似子序列碰撞,在精确计算前拒绝大多数碰撞对。Panache自行计算所有依赖数据的参数,只需调整资源预算。在17个UCR配置上,默认预算下它能恢复所有前20个泛基序,比本文基准测试的所有CPU和GPU方法都快。在有51个长度的500万个样本的晶圆数据集上,Panache在2.9分钟内完成一遍扫描,6.0分钟内发出精确基序,而最快的精确CPU基线需要7.95小时,H100 GPU上的SCAMP需要38.3分钟。

英文摘要

Motif discovery, the search for recurring patterns within a time series, is a core primitive of exploratory data analysis. A pattern, however, is defined by its duration, which analysts rarely know in advance. To resolve this unknown duration, an interval of window lengths is defined, and the accepted method is to try every length in that interval. Existing pan matrix profile (PMP) methods compute one z-normalized matrix profile per length, so $L$ lengths cost $L$ quadratic self-joins over the same series. We introduce Panache, to our knowledge the first one-pass streaming algorithm for z-normalized PMP motif discovery. It replaces the repeated self-joins with a single scan whose runtime is near-linear in the series length. The key observation is that mean-centering a subsequence changes only its DC Fourier coefficient, so the non-DC spectrum of every z-normalized subsequence can be maintained online by sliding-DFT recurrences and running statistics. This spectral state is the key under which similar subsequences collide in an occupancy-controlled hash directory and, through Parseval's theorem, yields a lower bound that rejects most colliding pairs before any exact computation. Panache computes every data-dependent parameter itself, leaving only a resource budget to tune. At the default budget, it recovers all top-20 pan-motifs against exact fixed-exclusion ground truth on 17 UCR configurations, and is faster than every CPU and GPU baseline benchmarked in this paper. On Wafer at five million samples over 51 lengths, Panache completes one pass in 2.9 minutes and emits the exact motifs in 6.0 minutes, against 7.95 hours for the fastest exact CPU baseline and 38.3 minutes for SCAMP on an H100 GPU.

Comments23 pages, 4 figures, 19 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑