arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.06509stat.ME

重新审视多重检验中的依赖性:用于FDP控制的经验分布方法

Revisiting dependence in multiple testing: empirical distribution approaches for FDP control

Fangyong Zheng, Pengfei Li, Yuan Jiang, Tao Yu

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出基于经验累积分布函数的eFDP方法,无需显式建模依赖性即可实现最优FDP控制,模拟和实际数据验证其更准确、功效更高。

中文摘要 AI 辅助

大规模多重假设检验是高通量数据分析的核心,其中控制错误发现至关重要。经典方法通常依赖于理论零分布,并常常针对检验统计量之间的依赖性进行调整,但当零统计量的经验分布偏离理论假设时,这些方法可能会产生误导。基于这一观察,我们研究了零检验统计量的经验累积分布函数(c.d.f.)在控制错误发现比例(FDP)中的作用。我们首先证明,在一种预言机场景下,即所有零假设的检验统计量的经验c.d.f.已知时,无论依赖结构如何,FDP控制都可以最优地实现,这凸显了对依赖性进行显式建模可能是不必要的。基于这一见解,我们提出了一种基于经验c.d.f.的FDP控制(eFDP)方法,通过多元混合模型框架和非参数估计程序来实现经验c.d.f.,建立了其渐近收敛性,并构建了一个实现渐近FDP控制的FDP控制程序。大量模拟表明,eFDP在强依赖性下尤其能比现有方法实现更准确的FDP控制和更高的功效,对高维乳腺癌基因表达数据集的分析证实了其实用性。

英文摘要

Large-scale multiple hypothesis testing is central to the analysis of high-throughput data, where controlling false discoveries is critical. Classical procedures typically rely on theoretical null distributions and often adjust for dependence among test statistics, but these approaches may be misleading when the empirical distribution of null statistics deviates from theoretical assumptions. Motivated by this observation, we investigate the role of the empirical cumulative distribution function (c.d.f.) of null test statistics in controlling the false discovery proportion (FDP). We first show that, under an oracle scenario where the empirical c.d.f. of the test statistics for all null hypotheses is known, FDP control can be achieved optimally regardless of the dependence structure, highlighting that explicit modeling of dependence may be unnecessary. Building on this insight, we propose an empirical c.d.f.-based FDP control (eFDP) method, implemented via a multivariate mixture model framework and a nonparametric estimation procedure for the empirical c.d.f.s, establish its asymptotic convergence, and construct an FDP control procedure that achieves asymptotic FDP control. Extensive simulations demonstrate that eFDP attains more accurate FDP control and higher power than existing approaches, particularly under strong dependence, and analysis of a high-dimensional breast cancer gene expression dataset confirms its practical utility.

发表机构

  • National University of Singapore(新加坡国立大学)
  • University of Waterloo(滑铁卢大学)
  • Oregon State University(俄勒冈州立大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑