arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

利用多结果进行样本高效CATE估计的表征学习

Representation Learning for Sample-Efficient CATE Estimation by Leveraging Multiple Outcomes

Maitreyi Swaroop, Shikha Bhat, Samantha Rodriguez, Tamar Krishnamurti, Bryan Wilder

arXiv 2609.06294首次发表:更新:

发表机构

Carnegie Mellon University; University of Pittsburgh(卡内基梅隆大学; 匹兹堡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出利用历史多结果数据学习低维表征,以在有限实验样本下高效估计CATE,理论证明可识别性并刻画偏差-方差权衡,实验验证其有效性。

AI 中文摘要

估计条件平均处理效应(CATE)能够实现干预措施的高效定向,但许多应用中的实验样本有限,这使得从高维协变量中估计异质性效应变得困难。在此类情境下,政策制定者和医疗从业者常常陷入维度灾难,或应用现成的降维方法,而这些方法可能无法保留处理异质性。然而,这些领域通常拥有大量历史数据集,测量了广泛的结果——这是一种在实践中很少被利用的监督来源。遵循因果表征学习,我们假设具有高维协变量的此类领域具有较低维度的潜在动态。因此,我们可以利用历史数据中测量的多样化结果来学习协变量的低维表征。理论上,我们证明当辅助结果满足一组替代条件且表征保留相关协变量信息时,当高维协变量被学习到的表征替换时,原始CATE是可识别的。结合现有的CATE估计的维度相关速率,该结果意味着在相同实验样本上具有更高的样本效率。此外,我们刻画了当假设不完全成立时的偏差-方差权衡,并表明当估计量方差的减少超过压缩带来的偏差时,基于表征的估计器仍能实现更低的误差。实证上,我们在合成数据和半合成医疗数据上评估了该方法。

英文摘要

Estimating conditional average treatment effects (CATE) enables efficient targeting of interventions, but many applications have limited experimental samples, making it difficult to estimate heterogeneous effects from high-dimensional covariates. In such settings, policymakers and medical practitioners often succumb to the curse of dimensionality or apply off-the-shelf dimension reduction methods that may not preserve treatment heterogeneity. Yet these domains often come with large historical datasets measuring a wide range of outcomes -- a source of supervision that is rarely exploited in practice. Following causal representation learning, we hypothesize that such domains with high-dimensional covariates have lower-dimensional underlying dynamics. We can thus leverage the diverse outcomes measured in historical data to learn a lower-dimensional representation of the covariates. Theoretically, we prove that when the auxiliary outcomes satisfy a set of surrogacy conditions and the representation retains relevant covariate information, the original CATE is identified when the high-dimensional covariates are replaced by the learned representation. Combined with existing dimension-dependent rates for CATE estimation, the result implies greater sample-efficiency on the same experimental sample. Additionally, we characterize the bias-variance tradeoff when the assumptions do not hold perfectly, and show that the representation-based estimator can still achieve lower error when the reduction in estimator variance outweighs the bias due to compression. Empirically, we evaluate the method on synthetic data and semi-synthetic medical data.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑