arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

实用的基于阈值的树编辑距离下界

Practical Threshold-based Tree Edit Distance Lower-Bounds

Lukáš Moravec, Radim Bača

arXiv 2609.03078首次发表:更新:

发表机构

VSB - Technical Universitty of Ostrava(奥斯特拉瓦理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对树编辑距离(TED)下界的权衡问题,通过实验比较现有方法,用Ukkonen算法加速字符串编辑距离(SED)下界,提出SED-struct阈值过滤器,在保持低成本的同时提升了异构树集合的过滤精度。

AI 中文摘要

对树结构数据使用树编辑距离(TED)进行基于阈值的相似度搜索计算量很大。给定一棵查询树和一个树数据库,目标是检索所有在预定义TED阈值τ内的树。由于精确的TED计算成本高昂,实用方法会使用下界在验证前修剪不相似的候选树。现有的下界存在一个基本权衡:低成本的统计和结构下界修剪能力有限,而更精确的基于遍历的字符串编辑距离(SED)下界使用标准二次动态规划计算成本高昂。此外,以往的比较研究未涵盖近期的结构过滤器或感知阈值的SED实现,导致它们的实用权衡尚不明确。本文首先在修剪精度和计算成本方面对最先进的TED下界进行全面实验比较;然后使用Ukkonen的有界字符串编辑距离算法加速SED下界,在不影响其修剪能力的情况下大幅降低运行时间;最后引入SED-struct阈值过滤器,该过滤器通过捕捉树节点间结构关系的轴感知约束来增强SED。在合成和真实世界数据集上的实验表明,SED-struct始终实现最高的过滤精度,同时保持实用的过滤成本。结果表明,SED-struct对于标准SED下界精度相对较低的异构树集合特别有益。

英文摘要

Threshold-based similarity search over tree-structured data using tree edit distance (TED) is computationally intensive. Given a query tree and a database of trees, the goal is to retrieve all trees within a predefined TED threshold $τ$. Because exact TED computation is expensive, practical methods employ lower-bounds to prune dissimilar candidates before verification. Existing lower-bounds exhibit a fundamental trade-off: inexpensive statistical and structural bounds provide limited pruning power, whereas the more precise traversal-based string edit distance (SED) bound is expensive to compute using standard quadratic dynamic programming. Moreover, previous comparative studies do not cover recent structural filters or threshold-aware SED implementations, leaving their practical trade-offs unclear. In this article, we first provide a comprehensive experimental comparison of state-of-the-art TED lower-bounds in terms of pruning precision and computational cost. We then accelerate the SED lower-bound using Ukkonen's bounded string edit distance algorithm, substantially reducing its runtime without affecting its pruning power. Finally, we introduce the SED-struct threshold filter, which strengthens SED with axes-aware constraints capturing structural relationships among tree nodes. Experiments on synthetic and real-world datasets show that SED-struct consistently achieves the highest filtering precision while retaining practical filtering costs. The results suggest that SED-struct is particularly beneficial for heterogeneous tree collections in which the standard SED lower-bound achieves relatively low precision.

CommentsFull version with 15 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑