基于训练子样本的模型辅助估计:一种结合设计方差估计的两阶段抽样方法
Model-assisted estimation with a training subsample: a two-phase sampling approach with design-based variance estimation
浏览论文内容
中文总结 AI 辅助
该研究针对训练子样本模型辅助估计的方差量化问题,将训练子样本视为第二阶段抽样,推导了方差分解与估计量,提出的方法可低成本恢复接近名义的覆盖率。
中文摘要 AI 辅助
当对从概率抽样中抽取的训练子样本拟合灵活预测模型时,实际报告的模型辅助估计量来自一个已实现的划分,但现有理论仅对其划分平均、交叉拟合或对称化版本量化不确定性。我们将训练子样本表示为第二阶段抽样,针对任意算法精确推导了两项方差分解及连接单划分估计量与其Rao-Blackwell平均的方差族,二者共享设计偏差。对于树型预测器,第二阶段方差可通过闭式计算,其在总方差中的占比随树复杂度增加而上升,解释了已记录的方差低估现象。我们提出了一种解析方差估计量和一种复制方差估计量,二者均不改变点估计,并通过模拟对其进行评估:分配第二阶段的成本可在仅为划分平均一小部分的代价下恢复接近名义的覆盖率。
英文摘要
When a flexible prediction model is fitted on a training subsample drawn from a probability sample, the model-assisted estimator actually reported arises from one realized partition, yet existing theory quantifies uncertainty only for partition-averaged, cross-fitted, or symmetrized versions of it. We represent the training subsample as a second phase of sampling and derive, exactly and for any algorithm, a two-term variance decomposition and the variance family linking the single-partition estimator to its Rao-Blackwellized average, whose design bias it shares. For tree-type predictors the second-phase variance is computable in closed form when the cell structure is fixed or s-measurable, and its share of total variance grows with tree complexity, contributing to documented variance underestimation through a mechanism distinct from residual shrinkage. We propose an analytic and a replication variance estimator, neither altering the point estimate, and evaluate them by simulation: in the populations studied, budgeting the second phase recovers most of the coverage lost by ignoring it, at a small fraction of the cost of partition averaging.
发表机构
- Universidad de la República(乌拉圭共和国大学)
机构由 AI 辅助整理,请以论文原文为准。