超额设计效应:聚类下阈值的有效样本量
The Exceedance Design Effect: Effective Sample Size for Thresholds under Clustering
浏览论文内容
中文总结 AI 辅助
该研究提出聚类下阈值的有效样本量的闭式定律,指出保形文献的校正量错误,数据集的有效样本量随阈值变化,偏差仅在单次部署时显现,并在25028样本集上验证约1300个阈值设置的可靠性。
中文摘要 AI 辅助
许多机器学习系统会在校准集的分位数处设置阈值:保形预测器通过将截止值设为校准集第90百分位来保证90%的覆盖率;弃权门控当模型得分低于校准集第10百分位时会拒绝回答;安全过滤器会屏蔽得分高于参考集第99百分位的任何输出。所有这些系统都承诺阈值在新数据上会保持规定的比率,该承诺假设校准样本相互独立,但在现代流程中,校准样本通常并非独立:它们共享同一个提示、文档或推理轨迹。自1965年以来,调查统计学已知如何通过计算一个样本相当于多少个独立观测值来对相关数据进行折扣,但该方法仅适用于平均值。我们表明,阈值需要不同的计数方式,该计数取决于聚类得分落在阈值同一侧的频率,且该频率随阈值设置位置的变化而变化,得分的数值相似度不影响该计数。我们证明了所得有效样本量以及部署系统实际观测到的覆盖率波动的闭式定律,由此得出三个结论:保形文献中当前使用的校正量是错误的,可能会向两个方向偏离;数据集没有单一的有效样本量,每个阈值设置水平对应一个有效样本量;这种偏差在多次运行的平均覆盖率中不可见,但会被单次部署的使用者完全感知。在一个包含25028个样本的公开校准集上,我们测量了约1300个(阈值设置下的)有效样本量对应的可靠性。
英文摘要
Suppose we want a cutoff that 90% of a population falls below. We estimate it from a sample, and another sample would give a different cutoff and a different fraction below it. We ask how much that fraction varies when observations come in independent groups, such as pupils in classrooms or sentences in news articles. We prove that grouping multiplies its large-sample variance by $1+(m-1)ρ_I(p)$, where $m$ is the group size, $p$ is the target fraction, and $ρ_I(p)$ measures whether two members of a group fall on the same side of the cutoff. That correlation can differ from the correlation between the scores themselves, and it changes with the target. We give a direct proof, a counterexample to using score correlation, and an extension to unequal group sizes. A dataset therefore does not have one effective sample size. How much information it contains depends on the question you ask. In our document experiment, the same 1,000 rows carried about 217 independent observations' worth of information at the median. At the 95th percentile, they carried about 621. Nothing about the dataset changed. We asked it a different question. The number of rows is a property of the dataset. The effective sample size belongs to the analysis.