发表机构
Massachusetts Institute of Technology(麻省理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对无黄金标准标签的AI生成数据的下游分析偏差问题,提出DMM框架,结合多个不完美AI测量实现有效推断,经模拟验证其有效性与效率提升作用,还开发了条件独立性假设的诊断方法。
AI 中文摘要
越来越多的学者使用AI来测量变量,随后将这些变量纳入下游分析。尽管AI测量的变量常被当作无误差观测值分析,但忽略自动测量中的预测误差会导致下游分析出现显著偏差和无效的置信区间,即便AI测量准确率很高(如超过90%)也是如此。现有解决方案如基于设计的监督学习和预测驱动推断,需将易出错的AI测量与黄金标准标签结合,而在某些应用领域,黄金标准标签可能成本高昂且难以获取。本文提出带多个不完美测量的去偏推断(DMM)框架,该框架结合多个易出错的AI测量,无需黄金标准标签即可实现有效的下游推断。基于CP分解的既定结果,DMM假设这些测量在潜在真实标签和观测单元级特征(如嵌入表示的文本特征)的条件下相互独立。该框架允许不同标注方法(如大语言模型)和不同标注单元(如文本)的未知误分类率存在差异。在此假设下,我们利用半参数推断理论证明DMM估计量是一致且渐近正态的,可支持社会科学中常见的各类下游统计分析的有效推断。模拟结果显示,DMM可实现有效推断,且添加准确但不完美的测量能提升效率。针对大语言模型标注的常见应用,我们还开发了诊断方法以评估条件独立性假设。
英文摘要
An increasing number of scholars use AI to measure variables they subsequently include in downstream analyses. Although AI-measured variables are often analyzed as if observed without error, ignoring prediction errors in automated measurement leads to substantial bias and invalid confidence intervals in downstream analyses, even if AI measurement accuracy is high, e.g., above 90%. Existing solutions, such as design-based supervised learning and prediction-powered inference, combine error-prone AI-based measurements with gold-standard labels, which may be costly and difficult to obtain in some application areas. In this paper, we propose debiased inference with multiple imperfect measurements (DMM), a framework that combines multiple error-prone AI measurements to enable valid downstream inference without gold-standard labels. Building on the established results on CP decomposition, DMM assumes that these measurements are independent conditional on the latent true label and observed unit-level features, such as text features represented by embeddings. This framework allows for unknown misclassification rates to vary across annotation methods (e.g., large language models) and across units of annotation (e.g., texts). Under this assumption, we use semiparametric inference theory to prove that the DMM estimator is consistent and asymptotically normal, enabling valid inference for a wide range of downstream statistical analyses common in the social sciences. Our simulation results show that DMM yields valid inference and that adding accurate, though imperfect, measurements can improve efficiency. Focusing on common applications of large language model annotations, we also develop diagnostics to assess the conditional independence assumption.