arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

小域中的SDR方差估计

SDR Variance Estimates in Small Domains

Eric Slud, Tim Trudell

arXiv 2608.17353首次发表:更新:

AI 中文总结

本文研究了小域中SDR方差估计的特性,明确了循环次数D的合理取值范围,通过理论与模拟揭示了SDR对小域平均方差的夸大效应及估计量的高变异性。

AI 中文摘要

连续差分复制(Successive Difference Replication,SDR)是Fay和Train(1995)提出的一种基于重复抽样的方差估计方法,适用于复杂多阶段调查的估计量,尤其是包含最终系统抽样阶段的调查。该方法多年来一直被用作美国人口普查局管理的大型全国家庭调查中的主要方差估计方法,包括美国社区调查(American Community Survey)以及基于自代表层的当前人口调查(Current Population Survey)月度估计。在其应用场景中,通常没有第二种方差估计方法可用,因此多位作者通过模拟研究了SDR在调查总量和非线性调查估计量方差方面的表现。本文首先详细阐述了SDR方法,并回顾了先前发表的关于SDR方差估计小域偏差的结果。研究表明,为了控制SDR估计量的变异性,实施SDR所用的循环次数D应不小于3,且无需大于5;除此之外,D的值对SDR小域偏差的发生几乎没有影响。通过理论公式和模拟显示,SDR方法会夸大小域中的平均估计方差,夸大幅度随连续枚举层中的属性均值、方差模式以及调查权重呈系统性变化。在样本量为20的域中,平均方差夸大程度通常为中等,不超过15%,但在特殊情况下可能更大。此外,SDR估计量在小域中具有极高的变异性,其标准差往往远大于任何偏差。

英文摘要

Successive Difference Replication (SDR) is a replication based method of variance estimation introduced by Fay and Train (1995) for estimators based on complex multistage surveys, especially those including a final systematic sampling stage. The method has been used for many years as the primary variance-estimation methodology in large national household surveys administered by the Census Bureau, including the American Community Survey and also the Current Population Survey's monthly estimates based on self-representing strata. In settings where it is applied, generally no second method of variance estimation has been available, so the performance of SDR has been studied via simulation by various authors, for variances of survey totals and of nonlinear survey estimators. This paper begins with a thorough exposition of the SDR method and review of previously published results on the small-domain biases of SDR variance estimation. It is shown that the number D of cycles used in implementing SDR should be 3 or larger, in order to control the variability of SDR estimates, but need not be larger than 5. Beyond that, the value of D is virtually irrelevant to the occurrence of small-domain bias in SDR. The SDR method is shown via theoretical formulas and simulation to inflate average estimated variances in small domains by amounts that vary systematically with the patterns of attribute means and variances and survey weights in consecutively enumerated strata. The degree of average variance inflation is generally moderate, no more than 15 percent in domains with sample size 20, but can be larger in special settings. Moreover, SDR estimates are extremely variable in small domains, with standard deviations often far larger than any biases.

Comments32 pages main text, 7 pages Supplement, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑