arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于图的概率链接:利用簇结构提升大规模实体解析的准确性

Probabilistic linkage with graphs: leveraging cluster structure to improve the accuracy of entity resolution at scale

Evelyn Lauren, Laura White, Trevor Graham Bell, Jacob Bor

arXiv 2610.00566首次发表:更新:

发表机构

University of the Witwatersrand; Boston University School of Public Health; National Institute for Communicable Diseases, National Health Laboratory Service; Africa Health Research Institute(威特沃特斯兰德大学; 波士顿大学公共卫生学院; 国家传染病研究所,国家卫生实验室服务; 非洲健康研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对大规模实体解析中传递闭包合并个体风险高的问题,评估了十一种聚类算法,并提出改编方案,发现Leiden、Louvain和标签传播等跨领域算法可减少高达66%的链接错误,且运行速度比马尔可夫聚类快40倍。

AI 中文摘要

概率记录链接即使在标识符不完整或不完美的情况下也能识别属于同一实体的记录,并越来越多地用于从大规模行政数据中进行人口层面的推断。链接会产生带有噪声的成对链接,这些链接随后必须被解析为不同的个体。标准方法——传递闭包——对每个链接独立地应用固定阈值,并从由此产生的连通分量中分配簇。然而,合并不同个体的风险随人口规模的增长而增加,这促使人们采用利用图结构来提高链接准确性的方法。相邻领域的聚类算法在其他领域解决了相同的图聚类任务,但从未被系统地评估用于实体解析。我们在三个模拟数据集和一个真实世界数据集上评估了十一种聚类算法——其中五种在实体解析领域已确立,六种来自相邻领域——并提出了一种针对链接图量身定制的改编方案。跨领域算法——Leiden、Louvain、标签传播——是最准确的算法之一,相对于传递闭包,将链接错误减少了高达66%。这三种算法运行速度也比马尔可夫聚类(实体解析文献中的领先方法)快多达40倍,并可直接集成到现有的大规模流程中。

英文摘要

Probabilistic record linkage identifies records belonging to the same entity even when identifiers are incomplete or imperfect, and is increasingly used to conduct population-level inferences from administrative data at scale. Linkage produces noisy pairwise links that must then be resolved into distinct individuals. The standard approach, transitive closure, applies a fixed threshold to each link independently and assigns clusters from the resulting connected components. However, the risk of merging distinct individuals grows with population size, motivating approaches that use graph structure to improve linkage accuracy. Clustering algorithms from adjacent fields solve the same graph-clustering task in other domains, yet have never been systematically evaluated for entity resolution. We evaluated eleven clustering algorithms -- five established in entity resolution and six from adjacent fields -- across three simulated datasets and one real-world dataset, and propose an adaptation tailored to linkage graphs. Cross-field algorithms -- Leiden, Louvain, label propagation -- were among the most accurate, reducing linkage error by up to 66% relative to transitive closure. The three algorithms also ran up to 40 times faster than Markov clustering (a leading approach in the entity resolution literature) and integrate directly into existing pipelines at scale.

CommentsCode will be made publicly available at https://github.com/evlauren/entity_resolution_clustering

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑