AI 中文总结
针对混合型变量关联度量的稳定性与计算效率问题,本文提出标签不变的总体度量ξ'及样本估计量ξ_n',经模拟和TCGA数据验证,该方法兼具编码稳定性、检验功效与计算优势。
AI 中文摘要
量化实值变量与分类变量之间的关联是数据分析中的基础任务。现有方法常依赖参数假设或任意整数编码,可能导致结果不稳定。我们提出一种针对实值-分类混合型场景的标签不变总体关联度量ξ',该度量归一化在0到1之间;当且仅当变量独立时取值为0,当且仅当分类变量是实值变量的可测函数时取值为1。我们还引入对应的样本估计量ξ_n',其计算复杂度为O(n log n)。这些度量对类别标签的排列和实值变量的严格单调变换具有不变性。我们证明了估计量ξ_n'的强一致性和渐近正态性,从而实现了无需排列的高效Wald独立性检验,以及总体度量ξ'的渐近置信区间。大量模拟实验和对癌症基因组图谱(TCGA)数据的应用表明,所提方法在名义混合型场景中具备编码稳定性、有竞争力的检验功效和显著的计算优势。
英文摘要
Quantifying the association between a real-valued variable and a categorical variable is a fundamental task in data analysis. Existing methods often rely on parametric assumptions or arbitrary integer encoding, which may lead to unstable results. We propose a label-invariant population measure of association, $ξ'$, specifically designed for the mixed real-valued-categorical setting. The proposed measure is normalized between 0 and 1; it equals 0 if and only if the variables are independent and 1 if and only if the categorical variable is a measurable function of the real-valued one. We also introduce a corresponding sample estimator, $ξ_n'$, computable in $O(n \log n)$ time. These measures are invariant to permutations of category labels and strictly monotone transformations of the real-valued variable. We establish the strong consistency and asymptotic normality of the estimator $ξ_n'$, enabling a computationally efficient, permutation-free Wald test for independence, and an asymptotic confidence interval for the population measure $ξ'$. Extensive simulations and an application to The Cancer Genome Atlas (TCGA) data demonstrate that the proposed method provides coding stability, competitive power, and substantial computational advantages in nominal mixed-type settings.
Comments82 pages, 17 figures. Submitted to the Electronic Journal of Statistics