arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29974cs.LG

多样几何,冻结权重:通过因果专家集成实现稳健的异质性处理效应估计

Diverse Geometries, Frozen Weights: Robust Heterogeneous Treatment-Effect Estimation via Causal Expert Ensembles

Ali Haghpanah Jahromi, Mohammad Taheri, Zohreh Azimifar

首次发表
浏览论文内容

中文总结 AI 辅助

针对观测数据中异质性处理效应估计的归纳偏置选择难题,提出五专家集成框架GeoACE,融合多种几何并冻结权重,在多个基准上取得稳健性能,验证了多样几何与无泄漏聚合的鲁棒性策略。

中文摘要 AI 辅助

从观测数据估计异质性处理效应是困难的,因为最合适的归纳偏置会随重叠程度、处理不平衡、预后结构和样本量而变化。我们引入了几何多样锚点校正专家集成(GeoACE),这是一个五专家框架,将共同的锚点校正估计器与互补的重叠感知和结果引导几何相结合。其任务级集成权重仅从内部验证预测中学习,在测试评估前冻结,然后应用于在完整开发样本上重新拟合的专家。第五个专家 O-Phi-ACE 从协变量和处理分配构建了一个无结果、重叠感知的统计投影,并用这个低维几何替换锚点输入。我们在八个基准协议上评估了 GeoACE 与 11 个比较器。加入 O-Phi-ACE 后,在所有七个具有个体效应真值的基准上,相对于四专家集成,平均 sqrt(PEHE) 有所降低,在 1,225 个配对任务中赢得了 998 个;在 JOBS 政策风险上的变化可以忽略不计。五专家集成在 IHDP100、IHDPA 和 IHDPB 上排名第一,在 NEWS 上排名第二,与 NEWS 领先者相差 0.13%。在七个 sqrt(PEHE) 基准中,它获得了最低的平均排名(3.714),尽管综合 Friedman 和 Iman-Davenport 检验不显著(p=0.328 和 p=0.330)。使用相同的五个冻结专家,在基准平衡分析中,逆倾向得分加权(inverse-DR weighting)始终优于赢家通吃选择、凸 DR 拟合、R 堆叠和因果 Q 聚合,但在统计上与等权重和 DR 岭收缩无法区分。因此,证据支持几何多样专家库和无泄漏聚合作为一种稳健性策略,而不是 GeoACE 或某种加权规则的普遍优越性。

英文摘要

Estimating heterogeneous treatment effects from observational data is difficult because the most appropriate inductive bias varies with overlap, treatment imbalance, prognostic structure, and sample size. We introduce the Geometry-Diverse Anchor-Correction Expert Ensemble (GeoACE), a five-expert framework that combines a common anchor-correction estimator with complementary overlap-aware and outcome-guided geometries. Its task-level ensemble weights are learned only from internal validation predictions, frozen before test evaluation, and then applied to experts refitted on the complete development sample. The fifth expert, O-Phi-ACE, constructs an outcome-free, overlap-aware statistical projection from covariates and treatment assignment and replaces the anchor input with this lower-dimensional geometry. We evaluate GeoACE against 11 comparators on eight benchmark protocols. Adding O-Phi-ACE reduced mean sqrt(PEHE) relative to the four-expert ensemble on all seven benchmarks with individual-effect truth, winning 998 of 1,225 paired tasks; the change on JOBS policy risk was negligible. The five-expert ensemble ranked first on IHDP100, IHDPA, and IHDPB and second on NEWS, differing from the NEWS leader by 0.13%. Across the seven sqrt(PEHE) benchmarks it obtained the lowest observed average rank (3.714), although the omnibus Friedman and Iman-Davenport tests were not significant (p=0.328 and p=0.330). Using the same five frozen experts, inverse-DR weighting was consistently better than winner-take-all selection, convex DR fitting, R-stacking, and causal Q-aggregation in benchmark-balanced analyses, but was statistically indistinguishable from equal weighting and DR ridge shrinkage. The evidence therefore supports geometry-diverse expert libraries and leakage-free aggregation as a robustness strategy, not universal superiority of either GeoACE or one weighting rule.

发表机构

  • University of Shiraz(设拉子大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑