高阶U中心化:ANOVA残差化与快速无偏估计
U-centering as subset ANOVA: edge regression and higher-order theory
AI总结:
该研究提出高阶U中心化方法,将其解释为ANOVA残差化,可快速无偏估计距离协方差等的高阶统计量,提升了计算效率并得到经典Hoeffding方差分量估计量。
AI中文摘要:
距离协方差和HSIC的无偏样本版本是四阶U统计量,而U中心化通过成对数组以$O(n^2)$运算对其进行评估。我们证明U中心化恰好是将加性端点效应拟合到对称空心数组后得到的最小二乘残差,该解释通过残差自由度解释了零行和及分母$n(n-3)$。我们将该构造扩展到由$r$子集索引的数组,高阶U中心化会移除所有涉及少于$r$个样本标签的效应,留下零$(r-1)$阶边际,并投影到维度为$\binom nr-\binom n{r-1}$的残差空间。对于具有$r$个参数的两个核,中心化数组的归一化内积对其第$r$个Hoeffding分量的配对是无偏的。尽管相应的直接估计量可涉及多达$2r$个不同观测值,但子集边际求逆或高阶U中心化后再进行归一化内积,可在固定$r$时以$O(n^r)$运算对其进行评估,相同计算还得到经典无偏Hoeffding方差分量估计量,其中最高分量表示为非负残差均方。
英文摘要:
The unbiased sample versions of squared distance covariance and the Hilbert-Schmidt independence criterion (HSIC) are fourth-order U-statistics, yet U-centering evaluates them from pairwise arrays in $O(n^2)$ operations. We show that U-centering is exactly the least-squares residual obtained after fitting additive endpoint effects to a symmetric hollow array. This interpretation explains the zero row sums and the denominator $n(n-3)$ through the residual degrees of freedom. The same pairwise residualization also gives useful regression identities. After endpoint effects are removed from both arrays, the U-centered dependence $t$-statistic is the ordinary slope $t$-statistic obtained by regressing one adjusted array on the other. In the two-sample problem, pooling the observations and using the between-group pair indicator as the predictor shows that the generalized-energy statistic is twice the fitted slope. The common-endpoint and fully interacted regressions give the same slope but use different residual standard errors. For $n\ge2r$, we extend the construction to arrays indexed by $r$-subsets. Higher-order U-centering removes all effects involving fewer than $r$ sample labels, leaves zero $(r-1)$-way margins, and projects onto a residual space of dimension $\binom nr-\binom n{r-1}$. For two symmetric kernels with $r$ arguments, the normalized inner product of the centered arrays is unbiased for the cross-moment of their $r$th Hoeffding components. A direct estimator can involve products spanning as many as $2r$ observations, but subset-margin inversion or higher-order U-centering evaluates the same quantity in $O(n^r)$ operations for fixed $r$. When both arrays are formed from the same kernel and sample, this becomes a nonnegative unbiased estimator of the variance of the highest-order Hoeffding component.