arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

划分差异:用于处理效应异质性与偏差的可解释因果森林

Splitting the Difference: Interpretable Causal Forests for Treatment Effect Heterogeneity and Bias

Nicolas Alexander Ihlo, Merle Behr

arXiv 2609.16971首次发表:更新:

发表机构

University of Regensburg(雷根斯堡大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出一种基于决策树和随机森林的简单算法,通过结合异质性与偏差校正两个分裂准则,无需额外估计倾向函数即可估计个体处理效应,兼顾预测准确性与可解释性。

AI 中文摘要

在医学和市场营销等多个领域,准确预测个体处理效应具有重大前景。然而,仅实现可靠的预测往往不足以做出明智的决策;同样重要的是理解为何某些个体的处理效应高于其他个体。为应对预测与解释这一双重挑战,我们提出了一种基于决策树和随机森林的算法,用于估计个体处理效应。我们的算法简单:其运行方式与标准随机森林完全一致,但采用不同的分裂准则,且无需额外的工作区,如广义随机森林中使用的双重机器学习或正交化。它能够处理具有不同处理倾向的观察性研究,而无需单独估计完整的倾向函数。这是通过结合两个分裂准则实现的——一个针对处理效应的异质性,另一个针对平均处理效应的偏差校正——两者共同改善分裂点选择,并自动将混杂因素与导致异质性的特征区分开来。因此,解释直接来自拟合的树结构本身,即树在哪些特征上分裂以及使用何种分裂统计量,而无需单独的事后分析。对于该算法的理论分析,我们考虑了一个带有潜在结果和处理倾向阶跃函数的变点模型,并提供了我们方法理论基础方面的见解。模拟研究表明,我们简单的算法在预测准确性上达到与现有方法相当甚至更好的水平,同时显著提高了可解释性。

英文摘要

In various fields, such as medicine and marketing, accurately predicting individual treatment effects holds significant promise. However, achieving reliable predictions alone is often insufficient for making informed decisions; it is equally important to understand why the treatment effect is higher for some individuals than for others. To address this two-fold challenge of prediction and interpretation, we introduce an algorithm based on decision trees and random forests for estimating individual treatment effects. Our algorithm is simple: it operates exactly like a standard random forest, but with a different splitting criterion, and requires no additional workarounds such as double machine learning or orthogonalization as used in Generalized random forests. It handles observational studies with varying treatment propensities without requiring separate estimation of the full propensity function. This is achieved by combining two splitting criteria---one targeting heterogeneity in the treatment effect, the other targeting bias correction for the average treatment effect---which together improve split point selection and automatically distinguish confounders from features responsible for heterogeneity. As a result, interpretation follows directly from the fitted tree structure itself, that is, from which features the trees split on and with which split statistics, without requiring separate post-hoc analysis. For the theoretical analysis of this algorithm, we consider a change point model with step functions for potential outcomes and treatment propensity and provide insights into the theoretical underpinnings of our approach. Simulation studies show that our simple algorithm achieves comparable, and often better, prediction accuracy than existing methods, while substantially improving interpretability.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑