arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

度量加权编辑距离:$\widetilde O_\varepsilon(N^{1.6})$ 时间内的 $(3+\varepsilon)$-近似算法

Metric Weighted Edit Distance: $(3+\varepsilon)$-Approximation in $\widetilde O_\varepsilon(N^{1.6})$ Time

Debarati Das, Evangelos Kipouridis, Tomasz Kociumaka

arXiv 2609.20796首次发表:更新:

发表机构

Max Planck Institute for Informatics(马克斯·普朗克 informatics 研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对度量代价的加权编辑距离,提出随机化 $(3+\varepsilon)$-近似算法,运行时间 $\widetilde{O}(N^{8/5})$,基于采样、廉价字符移除和平面图距离数据结构,并引入有界长度片段分解。

AI 中文摘要

对于每个 $0 < \varepsilon \le 1$,当代价在由字母表加上一个间隙符号构成的度量上形成时,我们给出加权编辑距离的随机化 $(3+\varepsilon)$-近似算法。对于总长度为 $N$ 的字符串,运行时间为 $\widetilde{O}(N^{8/5}/\varepsilon^{16/5})$,其中 $\widetilde{O}$ 抑制了 $\log(N/\varepsilon)$ 的多项式因子。对 $N$ 的依赖性与已知最快的单位代价编辑距离的 $(3+\varepsilon)$-近似算法相匹配。该算法从不低估编辑距离,并以 $N$ 的逆多项式失败概率实现近似保证。运行时间界限假设常数时间的精确算术运算和度量查询,并且与编辑代价的数值范围无关。我们基于三个工具:Chakraborty、Das、Goldenberg、Koucký 和 Saks(J. ACM, 2020)的采样框架,以及 Andoni(2020)的后续改进;Kuszmaul 对廉价字符的移除(ICALP 2019);以及 Klein 用于平面图中距离的数据结构(SODA 2005)。我们的新成分包括,除其他外,将一个字符串分解为具有高度结构化总删除代价的有界长度片段。这种分解使我们能够将所有片段与另一个字符串的一小族子串进行比较。

英文摘要

For every $0 < \varepsilon \le 1$, we give a randomized $(3+\varepsilon)$-approximation to weighted edit distance when the costs form a metric on the alphabet augmented with a gap symbol. For strings of total length $N$, the running time is $\widetilde{O}(N^{8/5}/\varepsilon^{16/5})$, where $\widetilde{O}$ suppresses factors polynomial in $\log(N/\varepsilon)$. The dependence on $N$ matches that of the fastest known $(3+\varepsilon)$-approximation for unit-cost edit distance. The algorithm never underestimates the edit distance and achieves the approximation guarantee with inverse-polynomial failure probability in $N$. The running time bound assumes constant-time exact arithmetic operations and metric queries, and it is independent of the numerical range of the edit costs. We build on three tools: the sampling framework of Chakraborty, Das, Goldenberg, Koucký, and Saks (J. ACM, 2020), with subsequent refinements by Andoni (2020); Kuszmaul's removal of inexpensive characters (ICALP 2019); and Klein's data structure for distances in planar graphs (SODA 2005). Our new ingredients include, among others, a decomposition of one string into pieces of bounded length with highly structured total deletion costs. This decomposition lets us compare all pieces against a small family of substrings of the other string.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑