arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

设计歧义感知的文书审查:一种用于记录链接与重复数据删除的分层抽样框架

Designing Ambiguity-Aware Clerical Review: A Stratified Sampling Framework for Record Linkage and Deduplication

Joseph Lam, Amaia Imaz Blanco, Efrosini Setakis, Jonny Laidler, Jonny Pearson, Katie Harron, Mario Cortina-Borja, Peter Christen, Giulia Mantovani

arXiv 2608.01401首次发表:更新:

AI 中文总结

该研究提出一种分层抽样框架,将文书审查视为有限总体抽样,用于记录链接与重复数据删除,通过权衡工作量、精度等提升评估效率,在预算约束下仍能保留高得分区间精度等关键信息。

AI 中文摘要

候选记录对的文书审查仍然是评估记录链接的事实上的黄金标准,但它资源密集且往往设计得不够正式。我们提出一种基于设计的框架,将文书审查视为对由匹配权重、比较模式、记录级歧义以及人口统计群组定义的细粒度层的有限总体抽样。匹配概率区间由基于模型的得分十分位数构建。在区间内,层将比较模式与由可匹配性和条件候选困惑度导出的歧义因子相结合。特定区间的误差边际轮廓编码了实质性优先事项,例如在高得分区间采用更严格的精度,而单个缩放参数则强制执行整体文书预算。我们使用在Splink中去重的带标签数据集评估该框架,该数据集包含50,000条记录和约478,000个候选对。我们将审查约23%对的基线设计与审查约7%对的预算约束设计进行比较。基线准确估计了全局和区间特定的匹配率,而比较模式、性别和歧义的抽样分布大致跟踪总体。在预算约束设计下,全局误差大约翻倍,最大的区间级误差出现在中间得分区间,其中匹配项、非匹配项和歧义案例相互混杂。最高得分区间的精度和区间级歧义轮廓在很大程度上得以保留,尽管比较模式和性别的代表性有所下降。该框架可推广到其他文书审查目标,并且可以将黄金标准数据作为设计和校准的先验信息纳入。它明确了工作量、精度、代表性以及链接不确定性覆盖范围之间可协商的权衡。

英文摘要

Clerical review of candidate record pairs remains the de facto gold standard for evaluating record linkage, but it is resource-intensive and often designed informally. We propose a design-based framework that treats clerical review as finite-population sampling over fine-grained strata defined by match weight, comparison pattern, record-level ambiguity, and demographic group. Match-probability bands are constructed from model-based score deciles. Within bands, strata combine comparison patterns with an ambiguity factor derived from matchability and conditional candidate perplexity. A band-specific margin-of-error profile encodes substantive priorities, such as tighter precision in high-score bands, while a single scaling parameter enforces the overall clerical budget. We evaluate the framework using a labelled dataset deduplicated in Splink, comprising 50,000 records and approximately 478,000 candidate pairs. We compare a baseline design reviewing about 23% of pairs with a budget-constrained design reviewing about 7%. The baseline accurately estimates global and band-specific match rates, while sampled distributions of comparison patterns, gender, and ambiguity broadly track the population. Under the budget design, global error approximately doubles, with the largest band-level errors in middle-score bands where matches, non-matches, and ambiguous cases are intermixed. Accuracy in the highest-score bands and the band-wise ambiguity profile are largely preserved, although representativeness by comparison pattern and gender declines. The framework generalises to other clerical-review objectives and can incorporate gold-standard data as prior information for design and calibration. It makes explicit the negotiable trade-offs between workload, precision, representativeness, and coverage of linkage uncertainty.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑