arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

噪声即信号:相关抽样误差对代理指标选择具有排序信息性

The Noise Is the Signal: Correlated Sampling Error Is Rank-Informative for Proxy Metric Selection

Sandro Provenzano

arXiv 2610.08194首次发表:更新:

发表机构

Zalando SE(Zalando SE)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文发现代理指标选择中共享抽样误差携带排序信息,去除它反而降低排序质量,并给出修正适用条件及廉价检查方法。

AI 中文摘要

客户终身价值等北极星指标通常过于缓慢且噪声大,难以用于短期A/B测试的决策。因此,团队依赖代理指标,通常根据其效果在过去实验中与北极星指标的追踪接近程度来选择。验证这一选择或任何做出该选择的方法都很困难:唯一的基准是带有噪声的北极星指标,且可用的历史实验数量有限。此外,代理指标和北极星指标的效果是在同一批客户上估计的,因此它们的抽样误差是相关的。主要实验平台近期的研究将这种共享误差视为污染予以去除,从而改进真实效应协方差的估计。然而,选择代理指标是一个排序问题,更好的估计未必带来更好的排序。我们通过在每个实验的客户不相交随机半样本上估计两种效应,来衡量无共享误差的一致性。在一个包含262个实验和69个候选代理指标的档案中,共享误差对候选指标的排序与这种一致性相似(Spearman相关系数为0.65):它携带了关于代理质量的信息。修正去除的共享误差越多,排序越差,因为去除丢弃了部分信号,但留下了排序噪声的主要来源——带有噪声的北极星指标和有限的实验数量——未受影响。档案校准的模拟(其中正确排序已知)证实了这一点,即使每次修正都获得真实的抽样协方差。在不相交客户半样本上评估的保留真实实验(使得共享误差无法使比较产生偏差)密切再现了预测的排序(Spearman相关系数为0.93)。随着实验数量的增加,修正仍可能带来收益,但所需数量随北极星指标的噪声急剧上升。我们绘制了这一交叉点,并为平台团队提供了三种廉价的检查方法,以根据其自身档案决定是否以及多大程度上进行修正。

英文摘要

North-star metrics such as customer lifetime value are often too slow and noisy to decide a short A/B test. Teams therefore rely on a proxy metric, commonly chosen by how closely its effects tracked the north star's across past experiments. Validating that choice, or any method for making it, is hard: the only benchmark is the noisy north star, and the number of available past experiments is limited. In addition, proxy and north-star effects are estimated on the same customers, so their sampling errors are correlated. Recent work at major experimentation platforms removes this shared error as contamination, improving estimates of the true-effect covariance. Choosing a proxy, however, is a ranking problem, and a better estimate need not give a better ranking. We measure agreement free of shared error by estimating the two effects on disjoint random halves of each experiment's customers. In an archive of 262 experiments and 69 candidate proxies, the shared error ranks the candidates in a similar order to this agreement (Spearman correlation 0.65): it carries information about proxy quality. The more of it a correction removes, the worse the ranking because removal discards part of the signal but leaves the main sources of ranking noise, the noisy north star and the limited number of experiments, untouched. Archive-calibrated simulations, in which the correct ranking is known, confirm this even when every correction receives the true sampling covariance. Held-out real experiments, evaluated on disjoint customer halves so that shared error cannot bias the comparison, closely reproduce the predicted ordering (Spearman correlation 0.93). Correction can still pay off with more experiments, but the number needed rises steeply with the north star's noise. We map this crossover and give platform teams three inexpensive checks for deciding from their own archive whether and how strongly to correct.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑