arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12489cs.LGstat.MEstat.ML

何时可信赖等开销Top-k分配的离线评估?一项受控、可复现的基准与从业者指南

When Can You Trust Offline Evaluation of Equal-Cost Top-k Allocation? A Controlled, Reproducible Benchmark and Practitioner's Guide

Binshuang Li

AI总结:

该研究构建等开销Top-k分配离线评估基准,分析弱重叠、优化器诅咒、倾向估计误差等问题,验证相关机制并发布仅含公开数据的可复现基准。

AI中文摘要:

机构在预算下决定治疗对象,希望在部署靶向规则前知晓其收益,离策略评估承诺从记录数据中获取该信息,但可部署规则是确定性Top-k策略,消除了所有动作平均,因此弱重叠会直接影响估计结果。我们在5个数据集和2个已知效应扫描中对6种估计器进行基准测试,并针对非模拟配对参考验证机制:其一,弱重叠由记录器-目标动作对齐而非仅记录器清晰度决定,支撑因素是记录器对目标动作的概率,从目标自身分数构建的记录器在测试范围内提升清晰度几乎不改变重叠,动作级不一致会使重叠崩溃;有效样本量可在不同记录环境中对该风险排序,但在从业者持有的单一日志内对候选者排序能力弱,且其阈值无法迁移。其二,优化器诅咒无法仅通过交叉拟合结果干扰项解决,当规则在用于评估它的数据上拟合时,仅交叉拟合干扰项会保留复用偏差并使其加剧,诚实策略级拆分通过针对学习过程的价值避免复用,这是估计量的改变而非全样本策略的去偏。其三,倾向估计误差是观测到的最大退化:折外估计对IPS的伤害大于施加的其他任何压力,使双重鲁棒估计几乎不变,还可反转重叠诊断;记录数据为合成生成,倾向值下限设为0.02,因此所有失效均发生在权重有界的情况下,下限还将两个调优混合器简化为未调优的父模型,留下4种实际不同的估计器,所有精确值表面均为合成或半合成,我们发布了仅含公开数据的基准。

英文摘要:

Organizations decide whom to treat under a budget and want to know what a targeting rule would have earned before deploying it. Off-policy evaluation promises this from logged data, but the deployable rule is a deterministic top-k policy: it removes all averaging over actions, so weak overlap hits the estimate directly. We benchmark six estimators across five datasets and two known-effect sweeps, and validate the mechanisms against a non-simulated paired reference. First, weak overlap is governed by logger-target action alignment, not by logging sharpness alone: what governs support is the logger's probability of the target's actions. Sharpening a logger built from the target's own score barely moves overlap over the tested range; action-level disagreement collapses it. Effective sample size ranks this risk across logging environments, but is weak at ranking candidates within the single log a practitioner holds, and its cut point does not transfer. Second, the optimizer's curse is not fixed by cross-fitting the outcome nuisance. When the rule is fit on the data used to evaluate it, cross-fitting the nuisance alone leaves the reuse bias in place and makes it worse. Honest policy-level splitting avoids the reuse by targeting the learning procedure's value -- a change of estimand, not a de-biasing of the full-sample policy. Third, propensity-estimation error is the largest degradation we measure: an out-of-fold estimate hurts IPS more than any other stress we apply, leaves doubly-robust estimation almost unchanged, and can invert the overlap diagnostic itself. Logging is synthesized and propensities floored at 0.02, so every failure occurs with bounded weights; the floor also reduces the two tuned hybrids to their untuned parents, leaving four practically distinct estimators, and all exact-value surfaces are synthetic or semi-synthetic. We release the benchmark; public data only.

↑