AI 中文总结
UpliftBench评估12种uplift估计器,发现基准分歧源于指标而非模型,揭示Qini等指标与效应准确性一致性不足等问题,发布可复现基准工具。
AI 中文摘要
Uplift建模(即条件平均处理效应估计)可用于个性化定向,但已发表的uplift基准测试常对哪种估计器表现最佳存在分歧;本文表明,这种分歧主要源于指标而非模型。UpliftBench采用外部测试隔离的多目标协议,在7个数据集系列上评估了12种uplift估计器;在存在参考目标的情况下,其两个核心发现为:标准连续基准IHDP上的F1指标,以及Jobs数据集内样本案例研究中的F2指标。在该基准上,Qini与效应准确性无显著一致性——在全部100个IHDP实现中,其与效应准确性的平均秩相关系数为+0.07,95%置信区间[-0.03, +0.16];而AUUC的一致性始终更高(配对前缀平均AUUC与Qini的差值为+0.49,95%置信区间[+0.40, +0.59];附带的累积增益AUUC一致性更高,为+0.73)。在Jobs数据集上,秩指标在符号阈值策略中存在结构性不足,因为它们丢弃了分数水平;经验上,在发布的拆分-旋转分析中,直接策略风险选择产生的基准遗憾低于随机模型选择,而Qini、AUUC和uplift-at-$k$则未做到这一点(遗憾为14-15%)。校准决策阈值可消除81%的Qini选择遗憾。两项发现均有边界而非普遍适用:在两个验证系列(ACIC和Revenue-Synthetic)中未检测到F1(两者差值均与零无显著差异),且在仅需秩的预算价值目标下F2消失。UpliftBench发布了带版本的加载器、固定协议、结果制品及可复现的动态排行榜;论文附带公开代码库。
英文摘要
Uplift modeling (conditional-average-treatment-effect estimation) drives personalized targeting, yet published uplift benchmarks frequently disagree on which estimator performs best; we show the disagreement is substantially about metrics, not models. UpliftBench evaluates 12 uplift estimators under an outer-test-isolated, multi-objective protocol across seven dataset families; its two findings are identified where a reference objective exists -- F1 on the standard continuous benchmark (IHDP), F2 in a within-sample case study on Jobs. On that benchmark, Qini shows no detectable alignment with effect accuracy -- across all 100 IHDP realizations its mean rank correlation with effect accuracy is +0.07, 95% CI [-0.03, +0.16] -- while AUUC is consistently more aligned (paired prefix-mean-AUUC-over-Qini gap +0.49 [+0.40, +0.59]; the shipped cumulative-gain AUUC aligns better still, +0.73). On Jobs, ranking metrics are structurally insufficient for a sign-threshold policy because they discard the score level; empirically, within the released split-rotation analysis direct policy-risk selection yields lower benchmark regret than random model selection while Qini, AUUC, and uplift-at-$k$ do not (14-15% regret). Calibrating the decision threshold removes 81% of the Qini-selection regret. Both findings are bounded, not universal: F1 is not detected on either validation family (the ACIC and Revenue-Synthetic gaps are both indistinguishable from zero), and F2 vanishes under a budgeted-value objective where rank suffices. UpliftBench releases versioned loaders, fixed protocols, result artifacts, and a reproducible living leaderboard; the public repository accompanies the paper.
Comments25 pages, 9 figures. Code, data, and results: https://github.com/binshuangli/uplift-bench