发表机构
Tsinghua University; Tencent Inc.(清华大学; 腾讯公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对CVR因果估计的样本选择偏差问题,提出带理论保障的双重鲁棒估计量,结合目标正则化提升稳定性,实验验证其优于简单去偏结合的方法。
AI 中文摘要
点击后转化率(CVR)是电商、广告等场景中的关键指标,反映转化流程第二阶段的效率与用户体验,因此估计CVR的因果效应具有重要实践意义。但直接将现有因果推断方法应用于点击样本,会因排除未点击数据而引入样本选择偏差并增大方差。近期CVR预测研究引入“理想损失”,利用全样本上损失的无偏估计优化模型参数,但无法保证损失的无偏性等价于最终估计量的无偏性。本文从半参数理论视角重新审视该挑战,针对CVR这类链式结构结果,提出一种新的双重鲁棒因果效应估计量,并详细推导其理论性质:该估计量相比干扰参数估计具有更快的收敛速率,因此在使用神经网络等灵活非参数估计器时更具鲁棒性。基于这些理论发现,本文进一步设计了基于目标正则化的框架,以提升数值稳定性与实际适用性。在合成数据和真实数据上的大量实验表明,所提方法的有效性与鲁棒性;此外,研究发现将损失去偏与标准因果估计量简单结合的效果不及本文方法,凸显了开发针对此类CVR类目标且具备坚实理论保证的新型估计量的必要性。
英文摘要
Post-click conversion rate (CVR) is a key metric in various scenarios including e-commerce and advertising, reflecting the efficiency and user experience in the second stage of the conversion process. Estimating the causal effect on CVR is therefore of great practical importance. However, directly applying existing causal inference methods to clicked samples introduces sample selection bias and increased variance due to the exclusion of non-click data. Recent studies on CVR prediction introduce "ideal loss", which optimizes model parameters using an unbiased estimate of the loss over the full sample. Nevertheless, there is no guarantee that unbiasedness of the loss implies unbiasedness of the final estimator. We revisit this challenge from the perspective of semiparametric theory. Specifically, we develop a new doubly robust causal effect estimator for chain-structured outcomes such as CVR, and derive its theoretical properties in detail. It achieves a faster convergence rate compared to nuisance parameters estimation and is therefore more robust when using flexible nonparametric estimators, including neural networks. Based on these theoretical findings, we further design a framework based on targeted regularization to improve numerical stability and practical applicability. Extensive experiments on synthetic and real-world data demonstrate the effectiveness and robustness of our method. In addition, we find that naively combining loss debiasing with standard causal estimators underperforms our method, highlighting the necessity of developing the new estimator tailored to this CVR-style objective with solid theoretical guarantees.