arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.40249cs.DS

低距离区域中的动态时间规整

Dynamic Time Warping in the Low-Distance Regime

  • Institute of Computer Science, University of Wrocław(弗罗茨瓦夫大学计算机科学研究所)
  • Department of Computer Science, Ariel University(阿里埃尔大学计算机科学系)
  • Max Planck Institute for Informatics(马克斯·普朗克信息学研究所)

机构由 AI 辅助整理,请以论文原文为准。

Itai Boneh, Shay Golan, Tomasz Kociumaka

AI总结:

本研究证明在低距离区域计算动态时间规整(DTW)在正交向量假设下需要 $n^{1-o(1)}k$ 时间,并给出在特定结构输入上的近线性算法,同时将下界扩展至模式匹配。

AI中文摘要:

动态时间规整(DTW)是一种经典的字符串和时间序列相似性度量,允许局部拉伸。给定字母表 $\Sigma$ 上的非空字符串 $S,T$ 以及代价函数 $\delta:\Sigma^2\to\mathbb{R}_{\ge0}$,$DTW_\delta(S,T)$ 是通过复制字符获得的 $S$ 和 $T$ 的等长扩展的最小总代价。对于长度至多 $n$ 的字符串,DTW 可在 $O(n^2)$ 时间内计算,并且在正交向量假设(OVH)下这是条件最优的。我们研究低距离区域,其中整数 $k$ 上界 $DTW_\delta(S,T)$,假设 $\delta(a,a)=0$ 且对于 $a\ne b$ 有 $\delta(a,b)\ge1$。对于几种经典相似性度量,该区域允许 $O(n+\operatorname{poly}(k))$ 算法,而对于具有度量代价的 DTW,已知最佳界为 $O(nk)$。我们证明这种依赖性本质上是最优的:假设 OVH,即使对于离散错配代价函数(每个错配分配代价 1),计算 DTW 也需要 $n^{1-o(1)}k$ 时间。下界适用于 $k$ 在常数和 $n$ 的线性之间的整个阈值谱。我们从正交向量的归约将向量坐标编码为相等字符游程的长度。产生的实例非常结构化:将游程折叠为单个字符揭示了具有短周期的长子串。我们补充了一个 $\tilde O(n+\operatorname{poly}(k))$ 时间算法,只要在输入中折叠游程后,每个周期为 $O(k)$ 的子串长度为 $\operatorname{poly}(k)$。最后,我们将此下界扩展到 DTW 模式匹配,该问题询问长度 $n$ 的文本的任何非空子串是否与长度 $m$ 的模式具有至多 $k$ 的 DTW 距离。我们证明经典的 $O(nm)$ 时间动态规划算法在 OVH 下是近最优的,即使当 $k=O(\log n)$ 时也是如此。

英文摘要:

Dynamic Time Warping (DTW) is a classical similarity measure for strings and time series that allows local stretching. Given non-empty strings $S,T$ over an alphabet $Σ$ and a cost function $δ:Σ^2\to\mathbb{R}_{\ge0}$, $DTW_δ(S,T)$ is the minimum total cost of equal-length expansions of $S$ and $T$ obtained by duplicating characters. For strings of length at most $n$, DTW is computable in $O(n^2)$ time, and this is conditionally optimal under the Orthogonal Vectors Hypothesis (OVH). We study the low-distance regime, where an integer $k$ upper-bounds $DTW_δ(S,T)$, assuming $δ(a,a)=0$ and $δ(a,b)\ge1$ for $a\ne b$. For several classical similarity measures, this regime admits $O(n+\operatorname{poly}(k))$ algorithms, whereas for DTW with metric costs the best known bound is $O(nk)$. We show that this dependence is essentially optimal: assuming OVH, computing DTW requires $n^{1-o(1)}k$ time even for the discrete mismatch-cost function, which assigns cost $1$ to every mismatch. The lower bound applies to the whole spectrum of thresholds $k$ between constant and linear in $n$. Our reduction from Orthogonal Vectors encodes vector coordinates in the lengths of equal-character runs. The resulting instances are very structured: collapsing runs to single characters reveals long substrings with short periods. We complement the lower bound with a $\tilde O(n+\operatorname{poly}(k))$-time algorithm whenever, after collapsing runs in the inputs, every substring with period $O(k)$ has length $\operatorname{poly}(k)$. Finally, we extend this lower bound to DTW pattern matching, which asks whether any non-empty substring of a length-$n$ text has DTW distance at most $k$ from a length-$m$ pattern. We prove that the classic $O(nm)$-time dynamic-programming algorithm is near-optimal under OVH, even when $k=O(\log n)$.

补充信息

↑