arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

证据类型竞争:干预数据何时能教会语言模型因果方向?

Evidence-Type Competition: When Can Interventional Data Teach Language Models Causal Direction?

Xining Xun

arXiv 2607.29484首次发表:更新:

发表机构

Tsingjiao Information Science (Beijing) Co., Ltd(北京清教信息科技有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究发现干预数据训练语言模型因果推理的金标准假设存在局限,观测上下文会抑制模型的因果方向判断能力,提出证据平均方案可降低符号错误率。

AI 中文摘要

干预数据被广泛视为训练模型因果推理的金标准。我们在完全受控的合成环境中测试这一假设,对比观测相关性与因果效应,发现其失效的情况颇具启发性。在辛普森悖论场景中,当两者符号系统相反时,增加预训练中干预样本的比例并不能提升因果方向的判断能力:模型的do()-响应幅度单调增长,但其符号却复制自观测上下文。决定是否使用干预证据的并非训练混合比例,而是推理时上下文存在的证据类型。在相同训练方案下,纯观测上下文在50个场景中引发29个符号反转,混合上下文引发19个,仅对齐干预探针则实现41个正确。从上下文中移除观测证据会立即释放被抑制的因果插值能力(true率提升+0.56);四状态内容操纵显示,该切换由内容介导且呈梯度变化。这种抑制在不同训练种子间稳定存在(匹配协议的第二个种子上11/11出现强反转),在0.93B参数规模下仍具鲁棒性(匹配探针仅组的反转率为31.8% vs. 6%),尽管绝对增益缩小至四分之一。对CLadder的外部审计揭示了学习到的正效应先验具有两层结构:符号随机重训练可在分布内移除该先验,但分布外无法移除。我们总结:能力存在于权重中,切换存在于上下文中,激活补丁将切换定位到中间层的观测行。我们进一步量化了基于探针的因果评估的采样噪声底限,以及将符号错误从26%降至9%的证据平均方案。

英文摘要

Interventional data is widely regarded as the gold standard for teaching models causal reasoning. We test this assumption in a fully controlled synthetic environment pitting observational correlation against causal effect, and find it fails instructively. In Simpson's-paradox worlds, where the two have systematically opposite signs, increasing the fraction of interventional samples in pretraining does not improve causal direction: the magnitude of the model's do()-response grows monotonically, yet its sign is copied from the observational context. What governs whether interventional evidence is used is not the training mixture but the evidence type present in the context at inference time. Under an identical training recipe, a purely observational context induces systematic sign reversal in 29/50 worlds, a mixed context in 19/50, while aligned interventional probes alone yield 41/50 correct. Erasing observational evidence from the context immediately releases the suppressed causal interpolation ability (ratio_true = +0.56); a four-state content manipulation shows the switch is content-mediated and graded. The suppression is stable across training seeds (11/11 strong reversals persist on a matched-protocol second seed) and robust as a rate at 0.93B parameters (31.8% vs. 6% reversals in the matched probe-only arm), even as absolute gains shrink four-fold. An external audit on CLadder exposes a learned positive-effect prior with a two-layer structure: sign-randomized retraining removes it in-distribution but not out-of-distribution. We summarize: the capability lives in the weights; the switch lives in the context, and activation patching localizes the switch to the middle layers' observational rows. We further quantify the sampling noise floor of probe-based causal evaluation and an evidence-averaging protocol that cuts sign errors from 26% to 9%.

Comments13 pages, 6 figures, 4 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑