arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.26298cs.LGcs.AI

实践中的实体解析:来自自助服务管道的经验教训

Entity Resolution in Practice: Lessons from a Self-Serve Pipeline

Kaushik Pavani, Ganga Aluri, Pravin Jadhav, Neeraj Prasad, Kiran Sanka

首次发表
浏览论文内容

中文总结 AI 辅助

本研究构建并评估了涵盖不同规模数据集的自助服务实体解析系统,得出3条现有文献未提及的实践经验,旨在帮助从业者减少无效实验。

中文摘要 AI 辅助

我们构建并评估了一个自助服务实体解析(Entity Resolution, ER)系统,该系统涵盖6个基准数据集,记录规模从864条到500万条不等,从中得出了现有ER文献中未提及的3条经验教训:(1)没有单一匹配算法在所有场景中都最优——由于自助服务管道无法预测下一个数据集,我们建议为每个数据集训练多个算法家族,并通过自动对比选出最优者;(2)精确率与召回率需分别优化,而非共享阈值——精确率需要基于硬规则的否决机制,召回率则需要更多样的候选检索;(3)一个误判的关联会悄然合并不相关实体——若假设“A匹配B”且“B匹配C”则“A匹配C”,单个错误关联会连锁合并数百条记录,因此每一次跨组合并都必须主动重新验证。我们希望这些经验能帮助从业者节省我们曾花费数月进行无效实验的时间。

英文摘要

We built and evaluated a self-serve entity resolution (ER) system on six benchmarks spanning 864 to 5M records, and three lessons emerged that are absent from existing ER literature. (1) No single matching algorithm wins everywhere - a self-serve pipeline cannot predict its next dataset, so we recommend training several algorithm families per dataset and letting an automatic bake-off pick the winner. (2) Precision and recall need separate fixes, not a shared threshold - precision needs hard rule-based vetoes, recall needs more diverse candidate retrieval. (3) One false-positive link can silently merge unrelated entities - assuming "A matches B" and "B matches C" implies "A matches C" lets a single bad link chain hundreds of records together, so every cross-group merge must be actively re-verified. We hope these lessons save practitioners the months of dead-end experiments that led us to them.

↑