arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.15416stat.MEstat.AP

深度域分层下基于设计的推断:印度第71轮全国抽样调查中的教学语言与私立机构选择

Design-Based Inference under Deep Domain Stratification: Language of Instruction and Private-Institution Choice in India's NSS 71st Round

Aksaj Goel, Abhishek Bhattacharjee

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对大型家庭调查多次细分后统计脆弱的问题,开发了可审计的基于设计的推断框架,结合印度第71轮全国抽样调查数据验证了该方法的有效性。

中文摘要 AI 辅助

大型家庭调查能够支持精确的全国估计,但在按地理、部门、性别、年龄和结果类别进行多次细分后,可能会在统计上变得脆弱。本文开发了一种可审计的基于设计的框架,用于确定分层多阶段调查中这种细分可以进行到何种程度。该框架围绕嵌套贡献账本构建,通过第一阶段与规模成比例的概率扩展、确定性加随机的村庄组选择以及第二阶段的家庭扩展来重建每个域的总和。非线性域参数表示为这些总和的比率,并通过一阶线性化进行分析。两个独立的全国抽样调查子样本提供了自然的重复方差估计量。粒度-稳定性轮廓将所得相对标准误与重复支持和集中度诊断相结合,使得详细估计能够附带关于设计是否能够支撑该估计的证据。本文确立了总和估计量的有限总体无偏性、比率线性化的渐近有效性以及线性化总和的两子样本方差估计量的无偏性。该方法通过第71轮社会消费:教育调查进行说明,重点关注印度和喜马偕尔邦的母语与教学语言,以及报告的偏好私立教育机构的原因。该应用保留了原始项目的实质性分析,同时将临时计算替换为可重复的推断工作流程。

英文摘要

Large household surveys support precise national estimates but can become statistically fragile after repeated disaggregation by geography, sector, sex, age, and outcome category. This paper develops an auditable design-based framework for deciding how far such disaggregation can be taken in a stratified multistage survey. The framework is built around a nested contribution ledger that reconstructs each domain total through the first stage probability proportional to size expansion, the certainty-plus-random hamlet-group selection, and the second-stage household expansion. Nonlinear domain parameters are expressed as ratios of these totals and analyzed by first-order linearization. The two independent National Sample Survey subsamples then provide a natural replication variance estimator. A granularity-stability profile combines the resulting relative standard error with replicate support and concentration diagnostics, so that a detailed estimate is accompanied by evidence about whether the design can sustain it. Finite population unbiasedness of the total estimator, asymptotic validity of the ratio linearization, and unbiasedness of the two-subsample variance estimator for linearized totals are established. The method is illustrated with the 71st-round Social Consumption: Education survey, focusing on home language versus medium of instruction and reported reasons for preferring private educational institutions in India and Himachal Pradesh. The application preserves the substantive analysis in the original project while replacing ad hoc calculation with a reproducible inferential workflow.

补充信息

↑