发表机构
Dell Technologies; University of Campinas(戴尔科技; 坎皮纳斯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一种广义U-统计量用于DNA序列进化模型的同质性检验,证明其渐近正态性并模拟评估有限样本性质。
AI 中文摘要
我们提出了一种检验统计量,用于在文献中一些最流行的进化过程下比较DNA序列。我们给出了该检验统计量的理论性质以及通过随机模拟得到的其经验表现。所提出的检验统计量是一种广义$U$-统计量,用于在零假设下检验分布同质性。我们表明这里存在一种二分情况。在零假设下,$U$-统计量的核是一阶退化的,该检验统计量属于拟$U$-统计量类,并遵循渐近正态分布,尽管其阶数高于标准情况。在异质性下,渐近正态性在更常见的一阶渐近性下获得。渐近正态性在以下情况下得到证明:高维/大样本量、高维/小样本量、低维/大样本量。此外,讨论了局部备择假设的情况,并建立了检验统计量对它们的邻接性。进行了模拟研究以评估检验统计量的一些有限维性质,涉及平衡/非平衡样本、维度和样本量等问题。
英文摘要
We present a test statistic for the comparison of DNA sequences under some of the most popular evolutionary processes available in the literature. Theoretical properties for the test statistic as well as its empirical performance by stochastic simulations are presented. The proposed test statistic is a generalized $U$-statistics built for tests of distributional homogeneity under the null hypothesis. We show that a dicothomous situation exists here. Under the null hypothesis, the $U$-statistics kernel is first-order degenerated, this test statistic falls in the quasi $U$-statistics class and follows an asymptotic normal law, albeit of higher order than the standard case. Under heterogeneity, the asymptotic normality is attained on the more usual first-order asymptotics. Asymptotic normality is proven for the cases: high-dimension/large sample size, high-dimension/small sample size, low-dimension/large sample size. Moreover, the case of local alternatives is discussed, and the contiguity of the test statistic for them is established. Simulation studies are performed to assess some finite-dimensional properties of the test statistic, regarding issues such as balanced/unbalanced samples, dimension and sample size.