AI 中文总结
本文提出干预分离选择(ISS)方法,通过查询干预来认证观测等价因果模型在干预下的一致性,并在连续变量上精确判定停止条件,实验证明其能有效审计网络因果抽象。
AI 中文摘要
观测等价的因果模型在干预下仍可能产生分歧,因为干预会创建在观测数据中从未出现的输入。我们引入了干预分离选择(ISS),该方法反复用幸存候选模型存在分歧的可允许干预查询真实系统,丢弃与结果相矛盾的候选模型,并在成本界限内没有干预能分离幸存者时停止。如果真实系统在候选模型中,这一停止条件保证每个幸存者与真实系统在界限内所有可允许干预上一致,这是任何观测学习器都无法提供的保证,无论其看到多少数据。停止条件仅依赖于幸存者,因此无需知道真实情况即可检查。对于连续变量,候选模型构成无限版本空间,混合整数线性规划可精确地在整个空间上判定停止条件,一致性在容忍度内成立。在一个三位彩色MNIST因果抽象任务中,墨水色调追踪数字大小,基于示例训练的普通卷积网络达到零留出错误,但在26%的单数字编辑上与基于形状的标签不一致,与基于色调的标签一样频繁。审计仅在此类图像上观测到的网络的因果抽象,ISS平均每张图像用13.6次干预认证每个网络的感知,每个证书针对所有可允许干预检查,只要网络的真实抽象在候选模型中即成立。当网络绕过每个候选抽象都依赖的单元时,覆盖该单元干预的证书可能被静默作废,二十次随机验证干预驳斥了其中69%的证书。
英文摘要
Observationally equivalent causal models can still disagree about what happens under intervention, because interventions create inputs that never occur in observational data. We introduce Interventional Separation Selection (ISS), which repeatedly queries the true system with an admissible intervention on which the surviving candidate models disagree, discards the candidates the outcome contradicts, and stops once no intervention within a cost bound separates the survivors. If the true system is among the candidates, this stopping condition certifies that every survivor agrees with it on every admissible intervention within the bound, a guarantee that no observational learner can give, however much data it sees. The stopping condition depends only on the survivors, so it can be checked without knowing the truth. For continuous variables the candidates form an infinite version space, and mixed-integer linear programs decide the stopping condition exactly over all of it, with agreement holding up to a tolerance. On a three-digit colored MNIST causal abstraction task in which ink hue tracks digit size, plain convolutional networks trained on examples reach zero held-out error, yet disagree with shape-based labels on 26% of single-digit edits, as often as hue-based labels do. Auditing the causal abstractions of networks observed only on such images, ISS certifies what each network perceives with 13.6 interventions per image on average, and each certificate, checked against every admissible intervention, holds whenever the network's true abstraction is among the candidates. When a network bypasses a unit that every candidate abstraction relies on, certificates covering interventions on that unit can be silently void, and twenty random validation interventions refute 69% of them.