arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

患病率决定精度:检测器定义数据集中的隐性污染

Prevalence Determines Precision:Silent Contamination in Detector-Defined Datasets

Jia Huang, Yankai Wan, Yangjun Ou

arXiv 2609.11449首次发表:更新:

发表机构

Guanghua School of Management, Peking University; College of Artificial Intelligence, Jilin University; School of Mathematical Sciences, Peking University(北京大学光华管理学院; 吉林大学人工智能学院; 北京大学数学科学学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究通过贝叶斯定理证明检测器定义数据集的精度由真阳性患病率决定,并揭示隐性污染作为第二信号影响估计器,提出需重新审视此类数据集的构建与使用。

AI 中文摘要

许多机器学习数据集是通过在候选池上运行检测器、启发式方法或模型来构建的;被接受的条目成为标签。数据集的精度由每个池中真阳性患病率通过贝叶斯定理决定,而不仅仅取决于检测器质量。使用单一仪器和时期,我们持有一个检测器定义的事件数据集,以及一个独立的官方索引,该索引将每个检测到的条目标记为真实或幻影。一个检测器,三个池产生幻影率分别为81.7%、9.0%和0.0%。将精度从两个高比率池转移到低比率池预测为0.955,而实测为0.183,误差为+422%;贝叶斯表达式在3.3%以内预测所有三个值。检测到的响应曲线是真实事件和幻影分量的精确凸组合(残差1.1e-16),幻影数量超过真实事件473对308,因此污染是第二个具有检测器继承形状的信号,而非加性噪声。污染方向取决于估计器:在相同窗口上,一个统计量被稀释,另一个被膨胀,因为其分母也被污染。一种常见的归一化将估计器转换为比率均值,其期望值可能不存在;在相同的335个事件上,它返回0.40,而定义良好的估计器返回0.10。

英文摘要

Many ML datasets are constructed by running a detector, heuristic, or model over candidate pools; accepted items become labels. Dataset precision is then governed by true-positive prevalence in each pool via Bayes, not solely by detector quality. Using one instrument and period, we hold a detector-defined event dataset plus an independent official index labeling every detected item as real or phantom. One detector, three pools yield phantom rates 81.7%, 9.0%, and 0.0%. Transferring precision from the two high-rate pools to the low-rate pool predicts 0.955 versus measured 0.183, a +422% error; the Bayes expression predicts all three within 3.3%. The detected response curve is an exact convex combination of a true-event and a phantom component (residual 1.1e-16), with phantoms outnumbering true events 473 to 308, so contamination is a second signal with detector-inherited shape, not additive noise. Contamination direction depends on the estimator: on identical windows one statistic is diluted and another inflated because its denominator is also contaminated. A common normalization turns the estimator into a mean of ratios whose expectation need not exist; on the same 335 events it returns 0.40 where the well-defined estimator returns 0.10.

Comments9 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑