arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31467stat.MEstat.AP

一种基于似然的生物医学独立性检验系数:二项截断复合似然比

A likelihood-based coefficient for biomedical independence testing: the binomial-cut composite likelihood ratio

发表机构国家过敏和传染病研究所 · 美国国立卫生研究院
查看机构详情
  • National Institute of Allergy and Infectious Diseases(国家过敏和传染病研究所)
  • National Institutes of Health(美国国立卫生研究院)

机构由 AI 辅助整理,请以论文原文为准。

Jing Qin

首次发表
浏览论文内容

中文总结 AI 辅助

提出一种基于复合伯努利似然比的独立性检验系数xi_cut,能捕捉非单调依赖,在模拟和蛋白质组数据中优于传统秩相关系数。

中文摘要 AI 辅助

生物标志物研究中使用的标准相关性度量——Pearson相关系数r、Spearman秩相关系数rho、Kendall秩相关系数tau——取值在[-1,1]之间,其中0表示不存在线性或单调关联。但零值无法区分独立性与非单调相关性,因此该尺度无法表示生物标志物实践中常见的阈值效应、异方差性和尾部偏移。我们将独立性检验构建为复合伯努利似然比:在每个阈值t处,比较1(Y≤t)在给定X条件下的伯努利分布与边际分布,并在所有截断点上进行聚合。所得系数xi_cut取值在[0,1]之间,当且仅当X与Y独立时(在Y连续的情况下)取0,当且仅当Y是X的可测函数时取1。Fisher加权在二阶近似下由伯努利似然自然产生,且xi_cut等于X与1(Y≤t)之间阈值平均互信息的二倍,从而给出I(X;Y)的无分布下界。二阶展开可恢复Fisher加权的Dette-Siburg-Stoimenov度量,该度量在连续性条件下与Chatterjee秩相关一致。估计采用Nadaraya-Watson插件法,带宽通过网格最大值选取;推断采用精确置换检验。在生物标志物驱动的模拟中,T_cut在W形非单调和异方差备择假设下显著优于基于秩的系数。我们在衰老血浆蛋白质组数据集的西雅图队列(n=70,年龄21-88岁)上进行验证,筛选全部1,305种蛋白质与年龄的相关性:在Benjamini-Hochberg控制q<0.05下,T_cut在70种蛋白质上拒绝原假设,其中6种是Pearson、Spearman和Chatterjee在相同FDR水平下遗漏的。

英文摘要

The standard dependence summaries used in biomarker studies -- Pearson's r, Spearman's rho, Kendall's tau -- take values in [-1, 1] with 0 indicating no linear or monotone association. Zero does not distinguish independence from non-monotone dependence, so the scale cannot represent threshold effects, heteroscedasticity, and tail shifts common in biomarker practice. We formulate independence testing as a composite Bernoulli likelihood ratio: at each threshold t, comparing the Bernoulli laws of 1(Y <= t) conditionally on X versus marginally, aggregated over cut points. The resulting coefficient xi_cut lies on [0, 1] with 0 iff X and Y are independent (under continuity of Y) and 1 iff Y is a measurable function of X. Fisher weighting arises at second order from the Bernoulli likelihood, and xi_cut equals twice the threshold-averaged mutual information between X and 1(Y <= t), giving a distribution-free lower bound on I(X; Y). A second-order expansion recovers the Fisher-weighted Dette-Siburg-Stoimenov measure, which coincides under continuity with Chatterjee's rank correlation. Estimation uses a Nadaraya-Watson plug-in with a max-over-grid bandwidth; inference is by exact permutation. In biomarker-motivated simulations T_cut substantially outperforms rank-based coefficients on W-shaped non-monotone and heteroscedastic alternatives. We illustrate on the Seattle cohort (n=70, ages 21-88) of the aging plasma proteome dataset, screening all 1,305 proteins for age dependence: under Benjamini-Hochberg control at q<0.05, T_cut rejects on 70 proteins, six of which are missed by Pearson, Spearman, and Chatterjee at the same FDR level.

补充信息

↑