arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.17084math.STstat.MEstat.TH

依赖条件下的镜像与仿冒+阈值

Mirror and knockoff+ thresholds under dependence

Xianyang Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

研究在缺乏特定性质下使用镜像与仿冒+阈值控制错误发现率(FDR)会出现的问题,通过构造满足特定条件的p值、考虑高斯零分数及不同尺度零分数等情况进行分析,揭示共享计数规则不能单独保证FDR,有效性取决于零符号联合行为。

中文摘要 AI 辅助

许多多重检验方法通过比较零分布的两侧来控制错误发现率(FDR)。小p值或大正分数被视为可能的发现,而大p值或大负分数用于估计这些发现中有多少是错误的。镜像和仿冒+阈值基于此想法构建。对于有效的仿冒统计量,这种比较由一个强性质证明是合理的:在它们的大小和非零分数的条件下,零符号是独立的公平硬币抛掷。本文探讨了在没有该性质的情况下使用相同阈值会出现什么问题。我们给出了三个答案。首先,我们构造了在一个子集上满足正回归依赖(PRDS)的精确均匀p值,其联合密度在单位立方体内处处为正。在10%的名义水平下,一个有11个假设的例子的FDR为17.4%,在同一族中FDR可能接近二分之一。其次,对于具有任何固定正等相关性的标准高斯零分数,无论多小,随着假设数量的增加,FDR最终会超过低于二分之一的每个水平。第三,当零分数可能具有不同尺度时,正定高斯模型可使FDR任意接近1。数值实验表明,这些失败在中等维度下是可见的。这些结果与仿冒理论并不矛盾。它们表明,共享计数规则本身并不能保证FDR:有效性取决于零符号的联合行为,而不仅仅取决于边际对称性、高斯性或PRDS。

英文摘要

Many multiple-testing procedures control the false discovery rate (FDR) by comparing the two tails of a null distribution. At a fixed cutoff, marginal symmetry makes this natural. Mirror and knockoff+ thresholds select the cutoff from the same data, so the standard finite-sample guarantee uses a stronger property: conditional on magnitudes and nonnull scores, null signs are independent fair coins. Failure can be severe without this property. Models satisfying positive regression dependence on a subset (PRDS) can have exactly uniform null $p$-values and large FDR. Under every fixed positive Gaussian equicorrelation, the all-null FDR converges to one half. Opposing loadings in Gaussian factor models can make FDR and power arbitrarily close to one; near-total failure also occurs for exchangeable, pairwise-uncorrelated scores. At a nominal input level $q<1/2$, no deterministic rule based only on the two tail counts can both reject and control FDR uniformly over our class if more discoveries or fewer controls cannot make rejection harder. We give finite-sample repairs based on joint sign information. Conditional sign odds may be bounded outside an exceptional event or averaged over negative controls; neither route uniformly dominates, and the integrated bounds are sharp. Independent calibration data or a specified Gaussian joint model yield valid adjusted levels. Simulations show that integration retains more power under diffuse Gaussian dependence, whereas exceptional-event calibration is more powerful when very large odds occur only for rare aligned signs; unadjusted FDR exceeds the target in both settings. Covariance alone is insufficient outside a specified joint model. Thus the relevant boundary is not marginal symmetry but joint information that remains valid after adaptive cutoff selection.

补充信息

↑