从因果合理性到因果可靠性:评估大型语言模型(LLM)作为校准的直接因果边分类器
From Causal Plausibility to Causal Reliability: Evaluating LLMs as Calibrated Direct Causal-Edge Classifiers
浏览论文内容
中文总结 AI 辅助
本研究评估12个指令微调开放权重LLM作为直接因果边分类器的可靠性,发现其判断过度预测、置信度不可靠,仅跨提示/模型一致性更优,建议将其作为软因果先验来源而非因果结构直接证据。
中文摘要 AI 辅助
大型语言模型(LLM)正越来越多地被用于为结构因果发现提供先验因果知识,但它们的直接边判断及其置信度是否可信仍不明确。我们在6个基准因果图、5种提示策略和4种置信度来源(口头表述、基于logit的、跨提示一致性、跨模型一致性)上,系统评估了12个经过指令微调的开放权重模型。在仅使用语言的成对协议下,我们的评估得出三个关键发现:(i)基于LLM的因果判断以召回率为主导:模型预测的图过于密集,存在大量假正例边,而提示主要改变精确率-召回率的权衡,而非解决过度预测;模型规模的增益在最大的图上会减弱,且无法消除校准误差。(ii)LLM通常能捕捉因果相关性,但无法可靠识别直接性或方向:相对于已发表的参考图,模型将40.0%的间接边和36.0%的反向非边误分类为直接边,而其他非边的误分类率为28.2%;此外,这些假正例中有80.8%和84.6%获得了至少80%的口头表述置信度,显示出对结构错误预测的过度自信。(iii)传统置信度估计不可靠,而一致性是更有前景的信号:基于logit的置信度无论正确与否都常接近1.0,而跨提示和跨模型一致性实现了更好的平均校准和区分度,但经Holm校正后其优势无统计学意义;基准熟悉度审计进一步在5对模型-数据集中识别出潜在的熟悉度,均涉及AsiaM。总体而言,我们的结果表明,LLM更适合作为经外部验证的软因果先验的来源,而非因果结构的直接证据。
英文摘要
Large language models (LLMs) are increasingly used to provide prior causal knowledge for structural causal discovery, yet whether their direct-edge judgments and confidence can be trusted remains unclear. We systematically evaluate 12 instruction-tuned open-weight models across six benchmark causal graphs, five prompting strategies, and four confidence sources: verbalized, logit-based, cross-prompt agreement, and cross-model agreement. Under our language-only pairwise protocol, our evaluation yields three key findings. (i) LLM-based causal judgments are strongly recall-dominant: models predict overly dense graphs with many false-positive edges, while prompting mainly shifts the precision-recall trade-off rather than resolving overprediction. Gains from model scale diminish on the largest graphs and do not eliminate miscalibration. (ii) LLMs often capture causal relatedness without reliably identifying directness or orientation. Relative to published reference graphs, models misclassify 40.0% of indirect and 36.0% of reversed non-edges as direct edges, versus 28.2% of other non-edges. Moreover, 80.8% and 84.6% of these false positives receive verbalized confidence of at least 80%, revealing substantial overconfidence in structurally incorrect predictions. (iii) Conventional confidence estimates are unreliable, whereas agreement offers a more promising signal. Logit-based confidence frequently collapses near 1.0 regardless of correctness, while cross-prompt and cross-model agreement achieve better mean calibration and discrimination, though their advantages are not statistically significant after Holm correction. A benchmark-familiarity audit further identifies potential familiarity in five model-dataset pairs, all involving AsiaM. Overall, our results suggest LLMs are better viewed as sources of externally validated soft causal priors than as direct evidence of causal structure.
发表机构
- Texas A&M University-Corpus Christi(德克萨斯农工大学科珀斯克里斯蒂分校)
- BITS Pilani Goa(戈亚斯 pilani Bits 大学)
- Mississippi State University(密西西比州立大学)
- Northern Illinois University(北伊利诺伊大学)
- Texas Christian University(得克萨斯基督教大学)
机构由 AI 辅助整理,请以论文原文为准。