发表机构
Mofid Securities(莫菲德证券)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对网络分布式多重检验,提出非渐近方法消除赢家诅咒,在有限样本下控制FDR,并保持通信效率。
AI 中文摘要
分布式多重检验要求 $N$ 个站点在严格的通信预算下控制全局错误发现率(FDR)。Pournaderi 和 Xiang(2024)的贪心区间聚合算法在渐近意义上解决了该问题,但在有限样本下可能违反 $\mathrm{FDR}\le\alpha$。我们将此违规追溯至所选密度统计量中的赢家诅咒偏差,其精确阶为 $\Theta(m^{-1/4}\sqrt{\log m})$,在标准带宽 $\varepsilon\asymp m^{-1/2}$ 下,其中 $m$ 为网络中 p 值的总数。交叉拟合贪心聚合(CFGA)通过在每节点数据的一半上选择嵌套拒绝族并在另一半上评分来消除该诅咒,当每节点零假设比例已知时,实现有限样本 $\mathrm{FDR}\le\alpha$;一种膨胀变体以 $\eta=1/m$ 的渐消松弛覆盖插件设置。BONuS-GA 则通过掩蔽由计数敲除(knockoffs)校准的合成均匀零假设袋,使得每个 p 值同时用于选择与推断;一种每节点预算变体移除所有预言机输入,对任何数据无关的袋控制 $\mathrm{FDR}\le\alpha$。通过随机分割上的 e 值聚合 CFGA 折叠(e-CFGA)消除了 Bonferroni 因子并平均了分割随机性。所有变体保持 $O(\sqrt{m}\log m)$ 的通信预算,除 e-CFGA 的 $O(\bar{R}\log m)$ 报告轮次外;经验上,BONuS-GA 在中等到较大每节点样本下占优,而 CFGA 在自适应带宽下于小样本占优。
英文摘要
Distributed multiple testing asks $N$ sites, each holding p-values for its own hypotheses, to control the false discovery rate (FDR) of the discoveries made across the whole network while communicating only a small number of bits. Greedy interval aggregation (Pournaderi and Xiang, IEEE TSIPN, 2023) meets the communication budget but controls FDR only asymptotically, and its FDR can exceed the target at finite sample sizes; the same interval counts both select the rejection regions and calibrate the stopping rule, which biases the selected interval densities upward. We propose \emph{budgeted BONuS-GA}: each site mixes synthetic uniform p-values into its data, the center ranks candidate p-value intervals across sites using the pooled counts, and each site's false discoveries are estimated from its own synthetic counts under a per-site share of the level. The procedure needs no knowledge of the null proportions, uses every p-value for both selection and inference, and controls $FDR\leα$ at every sample size for any fixed assignment of hypotheses to sites, assuming only that the null p-values are independent uniforms, independent of the non-null p-values. It keeps the original $O(\sqrt m\log m)$ communication, $m$ being the total number of p-values, when $N=O(\sqrt m)$. We also give a sample-splitting alternative, cross-fit greedy aggregation, with finite-sample control under a random-site model, and we quantify the selection bias behind the original procedure's failure. In a monitoring-network simulation the budgeted procedure retains most of the power of its heuristic counterpart at moderate-to-large site sizes.