AI 中文总结
针对二分图实验中部分分配导致估计偏差的问题,提出无偏线性估计器EARL,通过曝光与分配重加权实现最小方差,模拟显示误差显著低于现有基线。
AI 中文摘要
二分图A/B测试是一种实验设计,其中处理在单位的一个集合上随机分配,而结果则在另一个集合上测量。例如,在线市场可能在一组随机选择的商品上测试新的定价算法,而感兴趣的结果(如购买满意度)则在顾客身上测量,每位顾客会与许多商品互动。现有的二分图实验分析方法假设每个随机化单位都被分配到处理组或对照组。在实践中,通常只有一部分单位参与:平台会限制推出风险、保留留出组,并在并发测试中划分其总体。忽略未分配的单位会导致估计偏差,而包含它们则需要谨慎处理。我们构建了一个针对部分分配的二分图实验的无偏线性估计器。观测值不仅需要按参与连接中处理连接(曝光)的(中心化)比例进行重加权,还需要按参与连接的数量进行重加权,每个连接按其参与倾向的倒数加权,以便连接覆盖良好的单位按比例承担更多权重。由此产生的估计器EARL(曝光与分配重加权线性)是无偏、一致且渐近正态的;我们设计了两种渐近方差估计器,并证明它在自然类别的线性估计器中具有最小方差。无论未分配单位接受何种经验,只要它们对期望结果有加性贡献,EARL都保持无偏;特别是,它们可以被分配到其他不重叠的测试中。我们的理论辅以在两个公开数据集上的模拟研究,其中EARL的误差比最强现有基线低至六倍,在某些配置下,比Horvitz-Thompson式替代方案低一个数量级以上。
英文摘要
Bipartite A/B tests are experiments in which treatment is randomised over one set of units, while outcomes are measured on another. For example, an online marketplace may test a new pricing algorithm on a random subset of items, while the outcome of interest, say, purchase satisfaction, is measured on customers, each of whom interacts with many items. Existing methods for analysing bipartite experiments assume that every randomisation unit is assigned to treatment or control. In practice, often only a subset participates: platforms cap rollout risk, reserve holdout groups, and split their population across concurrent tests. Ignoring the unassigned units biases estimation, while including them requires care. We construct an unbiased linear estimator for bipartite experiments with partial assignment. Observations must be reweighted not only by the (centred) share of treated connections among participating ones (the exposure) but also by the number of participating connections, each inverse-weighted by its participation propensity, so that units whose connections are well covered carry proportionally more weight. The resulting estimator, EARL (Exposure- and Allocation-Reweighted Linear), is unbiased, consistent, and asymptotically normal; we devise two asymptotic variance estimators and show that it has minimal variance in a natural class of linear estimators. EARL remains unbiased regardless of the experience the unassigned units receive, as long as they contribute to expected outcomes additively; in particular, they can be allocated to other, non-overlapping tests. Our theory is complemented with a simulation study on two public datasets, in which EARL attains up to six times lower error than the strongest existing baseline and, in some configurations, over an order of magnitude lower than Horvitz-Thompson-style alternatives.
Comments16 pages, 2 figures