发表机构
VSB - Technical Universitty of Ostrava(奥斯特拉瓦理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对树编辑距离(TED)下界的权衡问题,通过实验比较现有方法,用Ukkonen算法加速字符串编辑距离(SED)下界,提出SED-struct阈值过滤器,在保持低成本的同时提升了异构树集合的过滤精度。
AI 中文摘要
对树结构数据使用树编辑距离(TED)进行基于阈值的相似度搜索计算量很大。给定一棵查询树和一个树数据库,目标是检索所有在预定义TED阈值τ内的树。由于精确的TED计算成本高昂,实用方法会使用下界在验证前修剪不相似的候选树。现有的下界存在一个基本权衡:低成本的统计和结构下界修剪能力有限,而更精确的基于遍历的字符串编辑距离(SED)下界使用标准二次动态规划计算成本高昂。此外,以往的比较研究未涵盖近期的结构过滤器或感知阈值的SED实现,导致它们的实用权衡尚不明确。本文首先在修剪精度和计算成本方面对最先进的TED下界进行全面实验比较;然后使用Ukkonen的有界字符串编辑距离算法加速SED下界,在不影响其修剪能力的情况下大幅降低运行时间;最后引入SED-struct阈值过滤器,该过滤器通过捕捉树节点间结构关系的轴感知约束来增强SED。在合成和真实世界数据集上的实验表明,SED-struct始终实现最高的过滤精度,同时保持实用的过滤成本。结果表明,SED-struct对于标准SED下界精度相对较低的异构树集合特别有益。
英文摘要
Threshold-based similarity search over tree-structured data using tree edit distance (TED) is computationally intensive. Given a query tree and a database of trees, the goal is to retrieve all trees within a predefined TED threshold $τ$. Because exact TED computation is expensive, practical methods employ lower-bounds to prune dissimilar candidates before verification. Existing lower-bounds exhibit a fundamental trade-off: inexpensive statistical and structural bounds provide limited pruning power, whereas the more precise traversal-based string edit distance (SED) bound is expensive to compute using standard quadratic dynamic programming. Moreover, previous comparative studies do not cover recent structural filters or threshold-aware SED implementations, leaving their practical trade-offs unclear. In this article, we first provide a comprehensive experimental comparison of state-of-the-art TED lower-bounds in terms of pruning precision and computational cost. We then accelerate the SED lower-bound using Ukkonen's bounded string edit distance algorithm, substantially reducing its runtime without affecting its pruning power. Finally, we introduce the SED-struct threshold filter, which strengthens SED with axes-aware constraints capturing structural relationships among tree nodes. Experiments on synthetic and real-world datasets show that SED-struct consistently achieves the highest filtering precision while retaining practical filtering costs. The results suggest that SED-struct is particularly beneficial for heterogeneous tree collections in which the standard SED lower-bound achieves relatively low precision.
CommentsFull version with 15 pages