发表机构
TripAdvisor, UK(TripAdvisor英国公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有树方法在CATE估计中的权衡痛点,提出融合显著性分裂与诚实样本拆分的混合算法,兼顾交互识别灵敏度与合法置信区间覆盖性能。
AI 中文摘要
估计异质性处理效应(CATE)需要同时完成效应修饰因子检测与估计不确定性量化。现有基于树的方法存在难以兼顾的权衡:基于显著性的方法(Radcliffe和Surry 2011)可直接识别子群交互作用,但无法提供有效推断;诚实因果树(Athey和Imbens 2016)可达到标称置信区间覆盖要求,但采用与结果无关的分裂准则,牺牲了交互作用检测灵敏度。本文提出一种混合算法,将基于显著性的分裂与诚实样本拆分、交叉验证相融合。所提分裂准则采用处理-协变量交互作用的t统计量平方$t^2$,当交互作用较强时,该准则被证明可与诚实$\text{EMSE}_\tau$准则直接对齐。事后诚实交叉验证用于选择代价复杂度惩罚项,得到在叶节点层面具备标称置信区间覆盖特性的单一原则性估计器。对于森林模型,本工作保留 bootstrap 计数向量以支持蒙特卡洛收敛的无穷小折刀法(IJ)方差估计,而非严格的逐点推断。在Athey和Imbens 2016提出的3种模拟设计上,单树模型在全部3种设计中均达到约90%的叶节点平均置信区间覆盖度,对应标称水平为90%(每种设计开展200次重复实验);在Criteo和Starbucks增益数据集上,所提方法的基尼系数性能与S-学习器、T-学习器基线持平。本工作配套开源Python包,包含可复现种子、兼容scikit-learn的API与完整测试覆盖,访问地址见对应链接。
英文摘要
Estimating heterogeneous treatment effects (CATE) requires simultaneously detecting effect modification and quantifying estimation uncertainty. Existing tree-based methods make an uneasy trade-off: significance-based approaches (Radcliffe and Surry 2011) identify subgroup interactions directly but lack valid inference; honest causal trees (Athey and Imbens 2016) deliver nominal confidence interval coverage but use outcome-agnostic splitting criteria that sacrifice interaction sensitivity. We introduce a hybrid algorithm that fuses significance-based splitting with honest sample-splitting and cross-validation. Our splitting criterion uses the squared $t$-statistic for the treatment $\times$ side interaction ($t^2$), which is shown to be directly aligned with the honest $\text{EMSE}_τ$ criterion when the interaction is strong. Post-hoc honest cross-validation selects the cost-complexity penalty, giving a single principled estimator with nominal CI coverage at the leaf level. For forests, we retain bootstrap count vectors to enable an infinitesimal jackknife (IJ) variance estimate of Monte-Carlo convergence rather than formal pointwise inference. On the three synthetic designs from (Athey and Imbens 2016) the single tree achieves approximately 90% leaf-average CI coverage at the 90% nominal level across all three designs (200 replications each); on the Criteo, Hillstrom and Starbucks uplift datasets we match Qini coefficient performance of S-, T-learner and GRF baselines. An open-source Python package with reproducible seeds, sklearn-compatible API, and full test coverage accompanies this work (https://codeberg.org/hadjipantelis/rattus).
CommentsTypos/omissions corrected, author name corrected