多阶段抽样设计下的极端总体选择及其应用
Extreme Population Selection under Multistage Sampling design With Applications
- Indian Institute of Technology Jodhpur(印度焦特普尔理工学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出两种序贯算法,在多阶段抽样下基于广义矩估计,无需参数假设即可高置信度识别极端总体,并在计量经济学和遗传学数据中验证了有效性。
AI中文摘要:
我们研究了在极端总体与最近总体充分分离的假设下,从K(≥2)个总体中选择极端(最优或最差)总体的问题。该选择基于一个适当的度量,该度量可能因应用领域而异。由于该度量的实际值未知,我们在多阶段抽样设计下使用广义矩方法获得其估计量。利用该估计量,我们提出了两种序贯程序,即在线算法和基于多臂老虎机的算法。在适当的正则性条件下,且不对底层分布施加参数假设,两种算法都能以期望的置信水平正确识别极端总体。我们通过计量经济学和遗传学中的应用来展示所提出的算法。在计量经济学应用中,使用基尼系数作为不平等度量来选择极端总体,并通过在各种分布设置下进行的广泛蒙特卡罗模拟研究来评估所提出程序的性能。在遗传学应用中,使用从肿瘤突变负荷(TMB)评分导出的度量来识别最差总体,并使用纪念斯隆凯特琳-IMPACT 50000临床测序队列证明了所提出算法的实际适用性。此外,我们使用所提出的框架来识别异常总体,前提是存在这样的总体。
英文摘要:
We study the problem of selecting the extreme (best or worst) population from among $K(\geq 2)$ populations, under the assumption that the extreme population is sufficiently separated from the nearest population. The selection is based on an appropriate measure, which may vary across different application domains. Since the actual value of the measure is unknown, we obtain its estimator using the generalized method of moments under a multistage sampling design. Using this estimator, we propose two sequential procedures, namely online algorithm and multi armed bandit based algorithm. Under suitable regularity conditions and without imposing parametric assumptions on the underlying distributions, both algorithms correctly identify the extreme population with a desired level of confidence. We illustrate the proposed algorithms through applications in econometrics and genetics. In the econometric application, the extreme population is selected using the Gini index as a measure of inequality and the performance of the proposed procedures is assessed through extensive Monte Carlo simulation studies conducted under various distributional settings. In the genetics application, the worst population is identified using a measure derived from the tumor mutation burden (TMB) score and the practical applicability of the proposed algorithms is demonstrated using the Memorial Sloan Kettering-IMPACT 50000 clinical sequencing cohort. Further, we use the proposed framework to identify an anomalous population, provided such a population exists.