AI 中文总结
本文提出一种在固定时间范围内对未知数量错误假设进行序贯检验的方法,通过先验加权混合和闭合检验实现速率最优的多重检验,并有效控制族系错误率。
AI 中文摘要
我们研究当观测成本高昂且预先指定采样时间范围时,对固定假设族进行序贯检验的问题。挑战在于,当错误假设的数量和身份未知时,如何汇集证据以支持早期决策,同时将任何错误拒绝的概率控制在水平 $α$ 下。现有的合并方法在单一错误假设数量下达到合并增长率:当只有一个错误假设时取平均,当全部错误时取乘积。在具有共同简单原假设和备择假设分布的独立流模型下,我们用边际似然比乘积的先验加权混合来检验每个交集。闭合检验结合这些初等对称多项式混合,以识别单个错误假设。特定设计的边界校准提供了有限时间范围内的族系错误控制,具有精确的有限状态保证或蒙特卡洛校准的置信度限定。先验匹配的混合在每个时间范围上唯一地最大化期望对数证据。对每个非空流子集赋予正权重的混合,当 $l$ 个流遵循备择假设时,达到对数增长率 $lD$。这里 $D$ 是每个备择观测的平均对数似然比,一轮为每个流提供一个观测。该速率在固定维度、配置、权重和足够长的时间范围内,当 $α\downarrow0$ 时达到一阶交集延迟下界。在给定截止日期下,单个假设的功效不能超过最佳单流功效,但当所有假设均为假时,闭合检验消除了多重性惩罚。高斯、篮子试验、语言模型和广告研究说明了这一点。主要篮子边界比 $1/α$ 低 40-52%。在 41 个模拟配置中,相对于预先指定的中期观察 Bonferroni 检验,封顶平均患者结局的中位减少率为 31%。
英文摘要
We study sequential testing of a fixed family of hypotheses when observations are costly and a sampling horizon is specified in advance. The challenge is to pool evidence for earlier decisions when the number and identities of false hypotheses are unknown, while controlling the probability of any false rejection at level $α$. Existing merges attain the pooled growth rate at a single number of false hypotheses: averaging when one is false, multiplying when all are. Under an independent-stream model with common simple null and alternative distributions, we test each intersection with a prior-weighted mixture of products of marginal likelihood ratios. Closed testing combines these elementary-symmetric-polynomial mixtures to identify individual false hypotheses. Design-specific boundary calibration gives finite-horizon family-wise error control, with exact finite-state guarantees or a confidence qualification for Monte Carlo calibration. The prior-matched mixture uniquely maximizes expected log evidence at each horizon. Mixtures assigning positive weight to every nonempty subset of streams attain log-growth rate $lD$ when $l$ streams follow the alternative. Here $D$ is the mean log likelihood ratio per alternative observation, and a round supplies one observation per stream. This rate attains the first-order intersection-delay lower bound as $α\downarrow0$ at fixed dimension, configuration, weights, and a long enough horizon. Power for an individual hypothesis cannot exceed the best single-stream power at a given deadline, but closure removes the multiplicity penalty when all are false. Gaussian, basket-trial, language-model and advertising studies illustrate both. The primary basket boundaries are 40-52% below $1/α$. Across 41 simulated configurations, the median reduction in capped mean patient outcomes relative to prespecified interim-look Bonferroni tests is 31%.