因果森林中的表示多重性
Representation Multiplicity in Causal Forests
浏览论文内容
中文总结 AI 辅助
本文研究因果森林中协变量的冗余编码(表示多重性)导致治疗决策改变及处理效应估计偏差,提出对生成相同分裂的变量组抽样以恢复预测不变性。
中文摘要 AI 辅助
在因果森林中,将协变量与严格单调的编码一同纳入,可能在不增加信息的情况下改变治疗决策。随机特征选择倾向于选择由多列表示的协变量。在给定条件下,我证明这种不平衡会随样本增长而持续存在。与其他协变量相关的处理效应分量,会根据其包含概率和树深度而被省略、削弱或恢复。模拟实验和一项职业培训的重复研究显示了对冗余编码的敏感性。当拟合和随机化保持不变时,对生成相同分裂的变量组进行抽样,可恢复分组数据上的预测不变性。
英文摘要
Including covariates alongside strictly monotone encodings can change a causal forest's treatment decisions without adding information. Random feature selection favors covariates represented by multiple columns. Under stated conditions, I show that this imbalance can persist as samples grow. Treatment effect components associated with other covariates are omitted, attenuated, or recovered depending on their inclusion probabilities and tree depth. Simulations and a job-training replication illustrate sensitivity to redundant encodings. Sampling groups of variables that generate identical splits restores prediction invariance on the grouping data when fitting and randomization are held fixed.
发表机构
- University of Pennsylvania(宾夕法尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。