发表机构
University of Warwick; University of Bergen(华威大学; 卑尔根大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对因果推理中结论迁移的问题,基于因果抽象理论提出模型级广义可迁移性框架,可同时处理所有目标查询,在近似场景下能给出有保证的查询区间,经实验验证有效。
AI 中文摘要
将因果结论从源研究总体迁移到目标总体是因果推理中的基础问题。可迁移性理论为此提供了判断标准:给定源总体的实验数据和目标总体的观察数据,该标准可判断目标查询是否可识别且判定过程是完备的,即若查询可迁移,该标准会给出精确公式。然而,该理论一次仅处理一个查询,返回的是表达式而非数值;且在两种具有实际重要性的场景下无法适用:查询不可迁移时,以及完全没有目标数据时。为解决这两个问题,我们基于因果抽象理论采取模型级视角:源和目标共享变量、图结构和干预方式,仅在已知的一组机制上存在差异,这使得可迁移性成为同层级抽象的特例。因此,我们不再询问单个查询是否可迁移,而是询问是否存在单一映射使源和目标的干预行为对齐。我们刻画了该映射在马尔可夫和半马尔可夫场景下的存在条件;当该映射存在时,所有目标查询可同时迁移。我们的主要贡献在于近似场景:当不存在精确映射时,最优近似映射仍可得到有保证的查询区间,将抽象误差重构为近似可迁移性的量化概念。我们将模型级迁移表述为对不可见目标的机制和环境扰动的分布鲁棒优化,并为两种具有挑战性的场景推导了保证:不可迁移查询的界,以及目标不可知场景下的保证。我们在合成马尔可夫和半马尔可夫基准数据集以及真实生态数据集上评估了我们的框架,结果显示有保证的区间能包含真实干预查询。
英文摘要
Transporting a causal conclusion from a source study population to a target one is a fundamental problem in causal inference. The theory of transportability provides a criterion for when this is possible: given experimental data from the source and observational data from the target, it determines whether a target query is identifiable and does so completely; i.e. if the query can be transported, the criterion finds the exact formula. However, it works one query at a time and returns an expression rather than the value itself. It is also silent in two practically important regimes: when the query is not transportable and when no target data exist at all. To tackle both, we take a model-level perspective grounded in Causal Abstraction theory. Source and target share variables, graph, and interventions, differing only at a known set of mechanisms, which makes transportability a special case of same-level abstraction. Thus, instead of asking whether one query transports, we ask whether a single map aligns the source and target across their interventional behaviour. We characterise when such a map exists in both the Markovian and semi-Markovian settings; when it does, every target query transports at once. Our main contribution lies in the approximate case. When no exact map exists, the best approximate one still yields certified query intervals, recasting abstraction error as a quantitative notion of approximate transportability. We formulate model-level transport as distributionally robust optimisation over mechanism and environment perturbations of the unseen target and derive certificates for both challenging regimes: bounds for non-transportable queries, and guarantees under target-agnostic settings. We evaluate our framework on synthetic Markovian and semi-Markovian benchmarks and a real ecological dataset, and we show that the certified intervals bracket the true interventional query.