arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

序贯超饱和筛选试验

Sequential Supersaturated Screening Experiments

Songqiao Han, Kalliopi Mylona, Steven Gilmour, Matteo Borrotti

arXiv 2609.37756首次发表:更新:

发表机构

King’s College London; University of Milano-Bicocca(伦敦国王学院; 米兰比可卡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对超饱和筛选试验中一次性方法选择因子过多、后续优化成本高的问题,提出两阶段序贯框架 S^3,通过正锥准则和分级分位数规则逐轮筛选,在模拟和 Borehole 基准上相比 LassoCV 降低了 I 类错误、提高了 F1 分数并加速优化。

AI 中文摘要

超饱和筛选试验用很少的试验次数研究许多候选因子。试验者必须首先识别活跃因子,然后优化响应。一次性正则化回归倾向于选择许多因子,而所选因子中的二阶响应面模型因此变得庞大,需要更多试验次数来拟合。我们提出了序贯超饱和筛选($S^3$),一个用于超饱和设计的两个阶段框架。$S^3$ 在一种新的正锥设计准则下,通过坐标交换构建每一轮筛选,该准则不使用从业者选择的 Welch 校准常数。然后,它使用分级分位数规则,每轮移除低重要性因子。该准则根据设计列相关性高于 Welch 下界的程度来评分,并适应当前的试验次数和候选因子数量。我们在每轮噪声因子存活和承诺的条件下,给出了噪声因子被固定的概率的逐轮界限,以及第一阶段总试验次数的显式上界。在模拟筛选问题和 Borehole 基准测试中,$S^3$ 相比一次性交叉验证 Lasso(LassoCV)实现了更低的 I 类错误和更高的 $F_1$ 分数。它在大多数设置中还提高了优化质量,并且运行更快。

英文摘要

Supersaturated screening experiments study many candidate factors with few runs. The experimenter must first identify the active factors, then optimize the response. One-shot regularized regression tends to select many factors, and the second-order response surface model in the selected factors is then large and needs many further runs to fit. We propose Sequential Supersaturated Screening ($S^3$), a two-stage framework for supersaturated designs. $S^3$ builds each screening round by coordinate exchange under a new positive-cone design criterion that uses no practitioner-chosen Welch calibration constant. It then removes low-importance factors one round at a time, using a graduated quantile rule. The criterion scores a design by how far its column correlations sit above the Welch lower bound, and it adapts to the current number of runs and candidate factors. We give a round-by-round bound on the probability that a noise factor is ever fixed, under conditions on noise survival and commitment at each round, together with an explicit upper bound on the total number of Stage 1 runs. Across simulated screening problems and the Borehole benchmark, $S^3$ achieves lower Type I error and higher $F_1$ scores than one-shot cross-validated Lasso (LassoCV). It also improves optimization quality in most settings and runs faster.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑