arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于高召回率候选生成的有限样本覆盖审核:认证与学习理论设计

Finite-Sample Coverage Audits for High-Recall Candidate Generation: Certification and Learning-Theoretic Design

Martin Anthony, Kaveh Salehzadeh Nobari

arXiv 2607.21480首次发表:更新:

发表机构

Data Science Institute, London School of Economics and Political Science; Department of Mathematics, London School of Economics and Political Science; The Inclusion Initiative, London School of Economics and Political Science(数据科学研究所,伦敦政治经济学院; 数学系,伦敦政治经济学院; 包容倡议组织,伦敦政治经济学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究高召回率候选生成中认证错过相关质量所需的审核标签数,刻画其标签复杂度,证明排除池审核的最优性,开发有限样本工具包,可认证错过质量、转换为召回率等,还能与审查负担配对选合适候选生成器。

AI 中文摘要

在经验管道的初始高召回阶段决定哪些项目进入后续审查、标记或建模,错过的相关项目会在后续阶段丢失。我们研究需要多少审核标签来以有限样本有效性认证错过的相关质量较小,主要结果刻画了该问题的标签复杂度。首先表明仅使用候选集内标签的程序无法认证错过质量的任何非平凡界限,审核必须对排除池采样。然后证明了匹配的有限语料库下限。基于此,开发了一个精确的有限样本工具包,可认证错过质量、通过双池设计转换为召回率、同时认证嵌套候选生成器的预指定族并生成针对声明扰动机制的压力测试证书。这些证书可与可观察的审查负担配对,以选择满足错过质量目标的负担最小的预指定候选生成器。每个保证都在一种规则下成立:候选生成器或从中选择它的预指定族以及审核规则在检查认证标签之前是固定的。

英文摘要

An initial high-recall stage in an empirical pipeline decides which items pass to later review, labelling, or modelling, and relevant items it misses are lost to every subsequent stage. We study how many audit labels are needed to certify, with finite-sample validity, that this missed relevant mass is small, and our main results characterise the label complexity of this problem. We first show that no procedure using only labels from inside the candidate set can certify any non-trivial bound on the missed mass: the audit must sample the excluded pool, the only region where unrecovered relevant items can lie. We then prove a matching finite-corpus lower bound. Any valid audit that certifies fewer than $m$ missed relevant items with high probability when none are present, even if adaptive and permitted to label the entire included pool, must inspect on the order of $N_0/m$ excluded-pool labels. Excluded-pool auditing is therefore minimax rate-optimal, not merely convenient, for missed-mass certification in the zero-miss regime. Building on this characterisation, we develop an exact finite-sample toolkit, using binomial and hypergeometric inversion rather than asymptotic approximation, that certifies missed mass, converts it to recall through a two-pool design, certifies pre-specified families of nested candidate generators simultaneously, and produces stress-test certificates against declared perturbation mechanisms. These certificates can be paired with observable review burden to select the least burdensome pre-specified candidate generator meeting a missed-mass target. Every guarantee holds under one discipline: the candidate generator, or the pre-specified family from which it is selected, and the audit rule are fixed before the certification labels are examined.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑