AI 中文总结
该研究分析潜在类别概率粗化为硬标签时的方向失真与覆盖率损失,推导插件估计量的收敛形式,指出偏差决定覆盖率,通过模拟和真实数据验证了argmax粗化的严重影响。
AI 中文摘要
研究将结果对未观测类别成员的校准概率向量进行回归。在排除得分的结构条件均值及条件校准下,观测数据模型可简化为部分线性回归。该概率向量是成员身份的Berkson型代理变量,因此效应向量τ可被识别且无衰减。实际应用中,该向量常被粗化为硬标签(如argmax、置信度阈值),且该标签无需保留Berkson性质。对于任意粗化,插件估计量收敛于𝒜τ,其中𝒜−I由得分被丢弃部分对保留部分的回归决定。因此,仅当该回归消失时,粗化才不会使τ失真,否则会使某些对比项的失真远大于其他项。同一算子决定了导致覆盖率损失的偏差。当该偏差与标准误差同阶时,Wald区间的极限覆盖率为Φ(z−ν)−Φ(−z−ν),其中ν为两者的比值;固定偏差会使覆盖率趋近于零。算子、标准误差以及通过未粗化估计量得到的偏差均可从观测数据中估计,因此可在报告区间前近似估算隐含的覆盖率。模拟实验显示,argmax粗化后会出现严重的覆盖率损失;三项真实数据检验则验证了方向特异性失真。
英文摘要
Outcomes are regressed on a calibrated probability vector for unobserved class membership. Under a structural conditional mean excluding the score and conditional calibration, the observed-data model reduces to a partially linear regression. The probability vector is a Berkson-type surrogate for membership, so the effect vector $τ$ is identified without attenuation. In practice the vector is often coarsened to a hard label - an argmax, a confidence threshold - and that label need not retain the Berkson property. For any coarsening the plug-in estimator converges to $\mathcal{A}τ$, where $\mathcal{A}-I$ is determined by the regression of the discarded part of the score on the retained part. Coarsening therefore leaves $τ$ undistorted exactly when that regression vanishes, and otherwise distorts some contrasts far more than others. The same operator determines the bias that drives coverage loss. Where that bias is of the order of the standard error, the Wald interval has limiting coverage $Φ(z-ν)-Φ(-z-ν)$, with $ν$ their ratio. A fixed bias sends coverage to zero. Operator, standard error, and - through the uncoarsened estimator - the bias are estimable from observed data, so the implied coverage can be approximated before the interval is reported. Simulations show severe coverage loss after argmax coarsening. Three real-data audits exhibit the direction-specific distortion.
Comments42 pages, 6 figures, 11 tables. Includes supplementary material with all proofs, two further identification results, additional experiments, and a land-cover audit