发表机构
Stanford University(斯坦福大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对聚类内依赖导致PPI置信区间覆盖率不足的问题,提出聚类稳健的PPI++,支持部分标注聚类,无需重抽样,并在电视新闻应用中实现95%覆盖率。
AI 中文摘要
数据收集往往成本高昂或后勤上要求苛刻,这限制了研究者能够探索的问题以及他们能够回答这些问题的精确程度。预测驱动推断(PPI)通过将标注数据与机器学习预测相结合,可以减少精确参数估计所需的数据量。然而,忽略聚类内部的依赖性可能导致置信区间覆盖真实参数的频率低于其名义水平。我们引入了聚类稳健的PPI++,它在独立聚类内任意依赖的情况下提供闭式标准误和渐近有效的置信区间,无需自助法或重抽样。我们的核心贡献是适应部分标注的聚类,这是一种常见的实证场景,其中聚类同时包含标注和未标注单元。由于单元在聚类内是依赖的,部分标注的聚类违反了PPI++的独立性假设。我们还展示了精度提升如何依赖于标注设计,并推导出一个聚类感知的功率调节规则,以最小化渐近方差。在电视新闻的应用中,标准PPI++置信区间的覆盖率低于60%,而聚类稳健的PPI++可以达到名义上的95%覆盖率。
英文摘要
Data collection is often costly or logistically demanding, limiting both the questions researchers can pursue and how precisely they can answer them. Prediction-powered inference (PPI) can reduce the amount of data needed for precise parameter estimation by combining labeled data with machine learning predictions. However, ignoring dependence within clusters can produce confidence intervals that cover the true parameter less often than their nominal rate. We introduce Cluster-Robust PPI++, which provides standard errors in closed form and asymptotically valid confidence intervals under arbitrary dependence within independent clusters, requiring no bootstrap or resampling. Our central contribution is to accommodate partially labeled clusters, a common empirical setting in which clusters contain both labeled and unlabeled units. As units are dependent within clusters, partially labeled clusters violate the independence assumption of PPI++. We also show how precision increases depend on the labeling design, and derive a cluster-aware power tuning rule that minimizes asymptotic variance. In an application to television news, standard PPI++ confidence intervals have coverage below 60%, whereas Cluster-Robust PPI++ can achieve nominal 95% coverage.