arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于快速大规模系统发育推理的自监督词汇表示学习

Self-Supervised Lexical Representation Learning for Fast, Large-Scale Phylogenetic Inference

Tim Wientzek

arXiv 2609.05262首次发表:更新:

AI 中文总结

本文提出一种自监督对比学习框架,从原始IPA词表学习词汇表示,无需同源性注释,可高效推断3399种语言的系统发育树,性能与基线相当,还能捕获概念稳定性,为大规模系统发育推理提供自动高效方案。

AI 中文摘要

计算系统发育学已成为历史语言学的重要工具,但其在全球规模的应用仍受两个因素限制:基于字符的方法所需的同源性判断的费力手动注释,以及对大型数据集进行推理的巨大计算成本。本文提出了一种完全自监督的对比学习框架,该框架直接从原始IPA转录的词表中学习词汇表示,无需同源性注释、对齐或额外的专家输入。该模型采用双重对比目标:将语音相似形式组织到连贯空间的词级损失,以及鼓励词汇空间反映语言更广泛语音特性的辅助语言级损失。从得到的词表示中,推导了成对语言距离,并用于推断包含3399种语言变体的全球系统发育树。该推断树与Glottolog参考树的广义四元组距离(GQD)与多个基线具有竞争力,而在标准笔记本GPU上仅需几分钟计算。此外,相同的表示捕获了历时概念稳定性:跨语言的成对距离方差产生的稳定性排名与已建立的排名显著相关。消融研究证实,语言级目标和语音特征向量的使用均改善了关于GQD的推断树拓扑。因此,该框架为大规模系统发育推理提供了一种计算高效且完全自动的替代方案,并提供了支持语言和概念层面下游分析的统一表示。

英文摘要

Computational phylogenetics has become an essential tool in historical linguistics, yet its application at a global scale remains constrained by two factors: the labor-intensive manual annotation of cognacy judgments required for character-based methods and the substantial computational cost of inference on large datasets. This paper introduces a fully self-supervised contrastive learning framework that learns lexical representations directly from raw IPA-transcribed wordlists, without requiring cognacy annotations, alignments, or additional expert input. The model employs a dual contrastive objective: a word-level loss that organizes phonetically similar forms into a coherent space, and an auxiliary language-level loss that encourages the lexical space to reflect broader phonological properties of languages. From the resulting word representations, pairwise language distances are derived and used to infer a global phylogenetic tree of 3,399 language varieties. The inferred tree achieves a generalized quartet distance (GQD) to the Glottolog reference tree competitive with multiple baselines, while requiring only minutes of computation on a standard notebook GPU. Furthermore, the same representations capture diachronic concept stability: variance in pairwise distances across languages yields stability rankings that correlate significantly with established rankings. Ablation studies confirm that both the language-level objective and the use of phonetic feature vectors improved the inferred trees topology with regards to GQD. The framework thus provides a computationally efficient and fully automatic alternative for large-scale phylogenetic inference and offers a unified representation supporting downstream analyses at both the language and concept level.

Comments27 pages, 3 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑