发表机构
Iowa State University(爱荷华州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对复杂调查设计下的半参数有效推断,提出统一理论框架,证明效率仅依赖极限包含概率,并给出交叉拟合估计量达到效率界限的条件,且通过模拟和实际数据验证。
AI 中文摘要
调查抽样固有的两个特征使半参数效率分析复杂化:抽样指标之间由设计引起的相依性,以及超总体规律下有限总体目标的随机性。对于一般的半参数全数据模型,我们证明在一大类相依设计下,观测数据试验是局部渐近正态的,其切空间结构与参考泊松试验相同。在联合超总体-设计规律下,一阶效率仅通过极限包含概率函数依赖于设计。参考试验中的标准随机缺失投影刻画了观测数据有效影响函数。有限总体目标通过一阶渐近展开处理,将分析从精确随机和扩展到非线性普查特征。超总体和有限总体中心化产生等价的局部正则性和效率概念,其界限通过一个给出广义有限总体修正的毕达哥拉斯分解相联系。然后,我们给出一般性和设计特定的条件,在这些条件下,具有估计 nuisance 函数的交叉拟合估计量能达到两个效率界限。对于有限总体均值,该界限等于 Godambe-Joshi 预期方差下界的大样本极限。对于标量目标,我们刻画了最优极限包含概率。模拟和加利福尼亚学术表现指数数据说明了该理论。
英文摘要
Two features intrinsic to survey sampling complicate semiparametric efficiency analysis: design-induced dependence among sampling indicators and the randomness of finite-population targets under the superpopulation law. For general semiparametric full-data models, we show that the observed-data experiment under a broad class of dependent designs is locally asymptotically normal with the tangent-space structure of a reference Poisson experiment. Under the joint superpopulation-design law, first-order efficiency depends on the design only through the limiting inclusion-probability function. Standard missing-at-random projection in the reference experiment characterizes the observed-data efficient influence function. Finite-population targets are treated through first-order asymptotic expansions, extending the analysis beyond exact random sums to nonlinear census characteristics. Superpopulation and finite-population centerings yield equivalent notions of local regularity and efficiency, with their bounds linked by a Pythagorean decomposition that gives a generalized finite-population correction. We then give general and design-specific conditions under which cross-fitted estimators with estimated nuisance functions attain both efficiency bounds. For the finite-population mean, the bound equals the large-sample limit of the Godambe-Joshi anticipated-variance lower bound. For scalar targets, we characterize optimal limiting inclusion probabilities. Simulations and California Academic Performance Index data illustrate the theory.