发表机构
Hong Kong University of Science and Technology; London School of Economics and Political Science; New York University; Rutgers University; Statistical Laboratory, University of Cambridge(香港科技大学; 伦敦政治经济学院; 纽约大学; 罗格斯大学; 剑桥大学统计实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究随机设计线性模型中回归 $F$-检验的双重稳健性,证明在误差分布或设计矩阵接近均匀时检验尺寸接近名义水平,并给出 Hölder 指数为 $1/3$ 或 $1/2$ 的量化结果。
AI 中文摘要
我们研究了随机设计线性模型中 $F$-检验的稳健性,并得出了一个较为微妙的结论。从积极的一面来看,我们的主要结果之一是:只要归一化误差向量的分布接近单位球面上的均匀分布,或者设计矩阵在经过适当的保持列空间的正交化方案处理后接近均匀分布,检验的尺寸就接近其名义水平。这在一定程度上表明 $F$-检验具有双重稳健性。我们通过建立 Kolmogorov 距离到 Wasserstein 距离的 Hölder 连续性性质来得出这一结论,该性质控制了原假设下 $F$-统计量与其名义 $F$-分布的偏离。设 $n$、$p$ 和 $p_0$ 分别为样本量以及全模型和零模型的维度,我们证明当 $\min(p-p_0,n-p) = 1$ 时 Hölder 指数为 $1/3$,当 $\min(p-p_0,n-p) \geq 2$ 时 Hölder 指数为 $1/2$。另一方面,这些指数相对较小,且一般情况下无法改进,这表明随着我们偏离检验精确成立的情形,检验的尺寸可能相当快地偏离其名义水平。在某些情况下,通过使用关注分布函数右尾差异的局部 Kolmogorov 距离,我们的结论可以得到改进。
英文摘要
We study the robustness of the $F$-test in random design linear models, and reach a somewhat nuanced conclusion. On the positive side, one of our main results is that the size of the test is close to its nominal level as soon as either the distribution of the normalised error vector is close to uniform on the unit sphere, or the design matrix, after applying a suitable column space-preserving orthogonalisation scheme, is close to being uniformly distributed. This provides a sense in which the $F$-test is doubly robust. Our conclusion is reached by establishing a Kolmogorov to Wasserstein distance Hölder continuity property controlling the departure of the $F$-statistic from its notional $F$-distribution under the null. Writing $n$, $p$ and $p_0$ for the sample size and the dimensions of the full and null models respectively, we prove that the Hölder exponent is $1/3$ when $\min(p-p_0,n-p) = 1$ and $1/2$ when $\min(p-p_0,n-p) \geq 2$. On the other hand, these exponents are relatively small and cannot be improved in general, revealing that the size of the test may depart from its nominal level quite quickly as we move away from settings where the test is exact. In some cases, our conclusions may be improved by working with a local Kolmogorov distance that focuses on discrepancies between distribution functions in the right tail.
Comments47 pages, 2 figures