外部化CPDAG摘要提升大语言模型因果推理能力
Externalized CPDAG Summaries Improve LLM Causal Deduction
浏览论文内容
中文总结 AI 辅助
针对因果推理中马尔可夫等价类易被思维链简化的问题,提出结构化思维方法,通过外部化受约束的CPDAG摘要提升LLM因果推理准确率,在Corr2Cause上显著优于基线。
中文摘要 AI 辅助
Corr2Cause任务询问一个因果主张是否在与观测到的相关性和条件独立性相容的每个有向无环图(DAG)中都成立。我们将此问题框架化为潜在对象推理:标签由CPDAG查询定义,但自由形式的思维链常常将马尔可夫等价类问题简化为局部模式匹配。我们提出结构化思维(Structured Thinking),这是一个两轮流水线,首先外部化一个类型化、受模式约束的CPDAG摘要,然后基于该图状态进行回答。在Corr2Cause完整测试集上,与强PC指令基线相比,结构化思维在主要配对运行中将Qwen3.5-27B的F1(是)分数从73.0提升至86.4(+13.4个百分点;McNemar检验p=2.4×10⁻⁶;bootstrap 95%置信区间[+8.4,+18.6]);在三个完整ID种子中,平均增益为+8.1±5.3个百分点。PC支架式两轮散文对照仅达到67.6的F1分数,表明详细的PC支架加上无模式散文中间步骤是不够的。相同模式在Qwen3.6-27B、Paraphrase-OOD和GPT-5.4-mini上保持一致。打乱生成的CPDAG导致F1分数下降12.0个百分点,完整分割审计显示与参考CPDAG高度一致(ID骨架F1为0.960;精确匹配率为75.9%)。这些结果支持一个有界设计原则:外部化定义标签的潜在对象,约束其形式,并测试下游答案是否使用它。
英文摘要
Corr2Cause asks whether a causal claim holds in every DAG compatible with observed correlations and conditional independencies. We frame this as latent-object reasoning: the label is defined by a CPDAG query, but free-form chain-of-thought often collapses the Markov-equivalence-class problem into local pattern matching. We propose Structured Thinking, a two-turn pipeline that first externalizes a typed, schema-constrained CPDAG summary and then answers against that graph state. On the Corr2Cause full test, Structured Thinking raises Qwen3.5-27B from $73.0$ to $86.4$ $F_1$(Yes) over a strong PC-instruction baseline in the primary paired run ($+13.4$ pp; McNemar $p=2.4\times 10^{-6}$; bootstrap $95\%$ CI [$+8.4$, $+18.6$]); across three full-ID seeds, the mean gain is $+8.1 \pm 5.3$ pp. A PC-scaffolded two-turn prose control reaches only $67.6$ $F_1$, indicating that a detailed PC scaffold plus a schema-free prose intermediate is not sufficient. The same pattern holds on Qwen3.6-27B, Paraphrase-OOD, and GPT-5.4-mini. Scrambling the emitted CPDAG costs $12.0$ pp $F_1$, and a full-split audit shows close agreement with the reference CPDAG (ID skeleton $F_1$ $0.960$; exact match $75.9\%$). These results support a bounded design principle: externalize the latent object that defines the label, constrain its form, and test whether downstream answers use it.
发表机构
- Nokia Bell Labs(诺基亚贝尔实验室)
- Télécom SudParis(巴黎南部电信学院)
- IRISA(法国国家信息与自动化研究所雷恩分部(注:IRISA常见中文名为雷恩信息与自动化系统研究所,为法国国家科研中心等共建的研究机构,通用译法为IRISA雷恩信息与自动化研究所))
机构由 AI 辅助整理,请以论文原文为准。