发表机构
University of Massachusetts Lowell(马萨诸塞大学洛厄尔分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对列联表效应量 Cramer's V 的无偏性不可能问题,提出将误差置于刀切法点估计和似然比投影区间,并给出引导规则与浏览器工具。
AI 中文摘要
Cramer's V 是几乎每个卡方检验旁都会报告的效应量,常与 Cohen 的标签对照解读,然而其中很大一部分数字描述的是抽样噪声,标准的偏差校正(Bergsma 2013)并不能解决这个问题。我们汇集了三个结果,并且只主张每个结果的推论。第一,平方效应量 phi^2 在任何样本量下都不存在无偏估计量:在固定 N 的多项抽样下,任何估计量的期望是单元概率的多项式,而 phi^2 不是。第二,无偏性仅在重标度是仿射变换时才随效应量的重标度转移,而在仿射选择中,V^2 = phi^2/k 是唯一有界于 [0,1] 且在完全关联时取值为 1 的。第三,我们给出了一个具有保守且渐近有效覆盖率的区间,该区间通过将单元概率的似然比置信集投影到效应量映射上获得;其下端点恰好当独立性检验未能拒绝时以闭式形式为零。因为没有任何东西是无偏的,唯一的问题是不可约误差置于何处。Bergsma 的校正将零误差置于原假设下,但在备择假设下误差为几个百分点;在 V^2 尺度上的删除一个的刀切法将误差薄薄地分散到各处(在 180 个设计网格上绝对偏差至多 0.008)。合并并不能消除偏差:在 3,000 次模拟元分析中,合并 200 项研究使朴素估计量落在真实值 0.01 以内的概率降至零,而刀切法的概率升至 0.99。投影区间在测试的设计上覆盖率为 0.997-1.000,而非中心反演在 phi^2 = 0.18 时覆盖率不足(0.936),且宽度为其两到三倍。点估计和区间是不同的问题,有不同的答案;将它们混为一谈正是文献中两者皆无的原因。我们给出了一个引导规则和一个实现该规则的浏览器工具。
英文摘要
Cramer's V is the effect size reported beside almost every chi-square test, read against Cohen's labels, yet a large fraction of such numbers describe sampling noise and the standard bias correction (Bergsma 2013) does not fix it. We assemble three results and claim only the consequence of each. First, the squared effect size phi^2 admits no unbiased estimator at any sample size: under fixed-N multinomial sampling the expectation of any estimator is a polynomial in the cell probabilities, while phi^2 is not. Second, unbiasedness transfers across a rescaling of the effect size only if the rescaling is affine, and among affine choices V^2 = phi^2/k is the one bounded in [0,1] with value 1 at perfect association. Third, we give an interval with conservative, asymptotically valid coverage, obtained by projecting a likelihood-ratio confidence set for the cell probabilities through the effect-size map; its lower endpoint is zero in closed form exactly when the test of independence fails to reject. Because nothing is unbiased, the only question is where the irreducible error is placed. Bergsma's correction puts zero error at the null but several percent under the alternative; a delete-one jackknife on the V^2 scale spreads it thin everywhere (absolute bias at most 0.008 over a 180-design grid). Pooling does not remove bias: across 3,000 simulated meta-analyses, pooling 200 studies drives the naive estimator's chance of landing within 0.01 of the truth to zero, while the jackknife's rises to 0.99. The projected interval covers 0.997-1.000 across the tested designs, while a noncentral inversion undercovers (0.936 at phi^2 = 0.18) at two to three times the width. The point estimate and the interval are different problems with different answers; conflating them is why the literature has neither. We give a routing rule and a browser tool that implements it.
Comments19 pages, 7 figures. Under review at Biometrical Journal. Reproducible code, data, and the derivations companion are openly archived on Zenodo: https://doi.org/10.5281/zenodo.21783660. Part of a contingency-table effect-size series (companions on the per-table noise floor, the exact conditional confidence interval, and the exact conditional distribution engine)