AI 中文总结
研究指出未观测混杂会使数据驱动因果发现中,增大样本量无法保证识别潜在因果结构,仅能提升观测代理变量关联的确定性,凸显统计不确定性降低与因果机制恢复的根本差异。
AI 中文摘要
未观测混杂是观察性因果推断公认的局限性,但其对数据驱动因果发现的影响常被低估,本文对此局限性给出渐近刻画。当真实因果变量未被观测但存在相关代理变量时,增大样本量并搜索日益稳定的数据驱动因果结构,会使涉及可观测代理变量的关联愈发确定。该结果凸显了减少统计不确定性与恢复因果机制的根本区别:更多数据可提升观测变量空间内真实关系的确定性,但无法保证识别出潜在因果结构。
英文摘要
Unmeasured confounding is widely recognized as a limitation of observational causal inference, but its implications for data-driven causal discovery are often understated. We provide an asymptotic characterization of this limitation. When the true causal variable is unobserved but correlated proxy variables are available, increasing sample size while searching for increasingly stable, data-driven causal structures will make associations involving observable proxies become increasingly certain. This result highlights a fundamental distinction between reducing statistical uncertainty and recovering causal mechanisms: more data can improve certainty about genuine relationships within the observed variable space while providing no guarantee of identifying the underlying causal structure.