arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

设计驱动的在线实验推断:证据收集、检验与估计的联合最优

Design-Driven Inference for Online Experiments: Jointly Optimal Evidence Collection, Testing, and Estimation

Jun Yu, Wenbiao, Zhao, Lixing, Zhu

arXiv 2610.07831首次发表:更新:

发表机构

School of Mathematics and Statistics, Beijing Institute of Technology; School of Science, China University of Mining and Technology-Beijing; Department of Statistics, Beijing Normal University at Zhuhai(北京理工大学数学与统计学院; 中国矿业大学(北京)理学院; 北京师范大学珠海校区统计学系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对在线多因素实验,提出设计驱动框架联合优化分配与证据规则,实现检验与估计双重最优,并通过模拟和LLM实验验证改进。

AI 中文摘要

在线多因素实验需要在连续监测下对实际可忽略的效应进行早期停止,同时控制错误停止,并保持对处理效应的精确估计。我们开发了一个设计驱动的框架,该框架联合选择分配和证据规则,将分配视为证据收集,将检验和估计视为同一证据的互补用途。对于具有干扰区组效应和处理-区组交互作用的多因素实验,我们证明了区组正交设计能从处理得分中去除干扰污染,并构成最坏情况下证据增长的完备类。在固定信息预算下,各向同性分配和显式径向e值联合达到方向未知备择假设的极小极大最优速率。同一分配对于处理效应的最大似然估计是普遍最优的,同时实现A-、D-和E-最优性,并确立了检验和估计的双重最优性。分批复制产生一个任意有效的e过程,保留累积条件最坏情况证据增长的极小极大最优速率,并保持预指定完整实验的普遍估计最优性。模拟和大型语言模型提示实验展示了停止效率和估计精度方面的改进。

英文摘要

Online multifactorial experiments requires early stopping for practically negligible effects while controlling false stopping under continuous monitoring and retaining precise treatment-effect estimation. We develop a design-driven framework that chooses allocation and the evidence rule jointly, treating allocation as evidence collection and testing and estimation as complementary uses of the same evidence. For multifactor experiments with nuisance block effects and treatment-by-block interactions, we show that block-orthogonal designs remove nuisance contamination from the treatment score and form a complete class for worst-case evidence growth. Under a fixed information budget, an isotropic allocation and an explicit radial e-value jointly attain the minimax rate optimal for directionally unknown alternatives. The same allocation is universally optimal for maximum likelihood estimation of treatment effects, simultaneously achieving A-, D-, and E-optimality and establishing double optimality for testing and estimation. Batchwise replication yields an anytime-valid e-process, retains minimax rate optimality for cumulative conditional worst-case evidence growth, and preserves universal estimation optimality for the prespecified complete experiment. Simulations and a large language model prompt experiment illustrate gains in stopping efficiency and estimation precision.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑