发表机构
Zhejiang University of Technology; Microsoft Research Asia; Sony Research; HKUST(浙江工业大学; 微软亚洲研究院; 索尼研究院; 香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文系统梳理监督式因果发现方法,区分可识别性与泛化性,强调模拟器假设影响预测,并提出匹配指标与测试分布变化的评估框架。
AI 中文摘要
监督式因果发现通过从带有结构标签的训练数据集对中学习,来推断新数据集的因果结构。这些训练对通常是模拟生成的,这使得模拟器既是监督的来源,也承载了关于因果图、机制和噪声的假设。因此,理解由此产生的预测结果,需要考察这些假设如何补充观测数据中可用的信息,而观测数据可能与多个因果图兼容。本文考察了截至2026年6月可用的代表性方法之间的这种关系。我们按预测目标、预测粒度、编码器、结构解码器和训练机制对这些方法进行组织,以将每种方法的预测内容与其使用数据和基于模拟器的监督方式联系起来。利用这一框架,我们区分了两个问题:目标在假设的模型类别下是否可识别,以及训练后的预测器是否能泛化到其训练分布之外。对机制和噪声的限制可以使原本模糊的因果方向变得可识别,但在这些限制下的预测准确性并不能在限制变化时建立迁移性。这一区分促使评估应使指标与可识别的图目标相匹配,并测试训练与部署之间图、机制和噪声的变化。将此类评估扩展到真实数据,还需要记录基准参考图背后的外部因果证据和不确定性。这些分析共同指导了方法比较,并指出了迁移、测试时适应和不确定性评估中的开放问题。
英文摘要
Supervised causal discovery learns to infer causal structure for a new dataset from training datasets paired with structural labels. These training pairs are typically simulated, making the simulator both a source of supervision and a carrier of assumptions about causal graphs, mechanisms, and noise. Understanding the resulting predictions therefore requires examining how these assumptions supplement the information available in observational data, which may be compatible with multiple causal graphs. This paper examines that relationship across representative methods available through June 2026. We organize these methods by prediction target, prediction granularity, encoder, structural decoder, and training regime to relate what each method predicts to how it uses data and simulator-based supervision. Using this framework, we distinguish two questions: whether the target is identifiable under the assumed model class, and whether a trained predictor generalizes beyond its training distribution. Restrictions on mechanisms and noise can make otherwise ambiguous causal directions identifiable, but predictive accuracy under those restrictions does not establish transfer when they change. This distinction motivates evaluation that matches metrics to the identifiable graph target and tests changes in graphs, mechanisms, and noise between training and deployment. Extending such evaluation to real data also requires documenting the external causal evidence and uncertainty behind benchmark reference graphs. Together, these analyses guide method comparison and identify open questions in transfer, test-time adaptation, and uncertainty assessment.