arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用图神经网络逼近系统发育树之间的SPR距离

Approximating SPR Distance Between Phylogenetic Trees with Graph Neural Networks

Renata Martins Castanheira, Miguel Bugalho, Cátia Vaz

arXiv 2607.18311首次发表:更新:

发表机构

ISEL -- Instituto Superior de Engenharia de Lisboa, Instituto Politécnico de Lisboa, Portugal(里斯本理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究用图神经网络逼近系统发育树SPR距离,构建公开数据集,建立预处理管道,验证监督目标,训练GIN回归器,其在分布内表现良好,外推有局限,发布的数据集等为SPR逼近提供可重复基础。

AI 中文摘要

比较系统发育树拓扑结构对于理解流行病动态至关重要,但像子树剪枝与重嫁接(SPR)距离这样具有生物学意义的距离计算是NP难的,在大型数据集上难以处理。我们研究图神经网络(GNN)能否在训练后每次比较时以近恒定时间逼近SPR距离。贡献有四点:一是构建并公开包含864个系统发育树及388个带标签树对的数据集;二是建立可重复预处理管道;三是验证监督目标;四是训练连体图同构网络(GIN)回归器,其在分布内表现良好,但外推到更大树时精度下降。发布的数据集及验证的启发式方法与精确关系为扩展学习的SPR逼近提供了可重复基础。

英文摘要

Comparing phylogenetic tree topologies is essential for understanding epidemic dynamics, yet biologically meaningful distances such as the Subtree Prune and Regraft (SPR) distance are NP-hard to compute and intractable on large datasets. We investigate whether a Graph Neural Network (GNN) can approximate SPR distances in near-constant time per comparison after training. Our contributions are fourfold. First, we build and publicly release a dataset of 864 phylogenetic trees inferred with UPGMA and Neighbor-Joining over four bacterial species, spanning up to 9{,}500 isolates, together with 388 labelled tree pairs. Second, we establish a reproducible pre-processing pipeline including midpoint re-rooting, which reduces tree depth and supplies the rooting required for exact distance computation and for the model's root-based features. Third, we validate the supervision target: on small trees, where exact SPR is tractable, the unrooted phangorn::SPR.dist heuristic correlates almost perfectly with the exact rooted distance computed by rspr (Pearson $0.98$--$0.99$), making it an excellent monotonic surrogate. Lastly, we train a Siamese Graph Isomorphism Network (GIN) regressor. In-distribution, i.e., held-out trees from the same species and size range as training, it explains roughly 87--90% of the variance ($R^2 \approx 0.87$ on a held-out split; $0.90 \pm 0.19$ under stratified cross-validation), with about four times lower error than a mean-predictor baseline, and shows partial transfer to unseen species ($R^2 \approx 0.37$). Its main limitation is extrapolation to trees larger than those seen in training, where accuracy collapses. The released dataset and the validated heuristic versus exact relationship provide a reproducible basis for scaling learned SPR approximation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑