arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

介入式基础审计:通过谓词替换对大语言模型思维链进行黑盒前提依赖性测试

Interventional Grounding Audits: Black-Box Premise-Dependency Tests for LLM Chain-of-Thought via Predicate Substitution

Hironao Nakamura

arXiv 2607.13069首次发表:更新:

AI 中文总结

研究大语言模型思维链前提依赖性问题,提出介入式基础审计方法,通过谓词替换干预前提并重新运行模型,在ProntoQA基准测试中检测前提依赖性表现优异,还发现了“正确答案,错误推理”信号。

AI 中文摘要

大语言模型产生的思维链推理看似逻辑合理,但可能并非真正依赖于所陈述的前提。我们引入了介入式基础审计,这是一种对前提依赖性进行黑盒、逐步骤的测试:通过用新符号替换单个前提的目标谓词来进行干预,重新运行模型,并检查每个推理步骤的标准化结论(规范谓词形式)是否改变。我们在具有金标准证明树的合成多跳演绎推理基准ProntoQA上进行评估,其中逐步骤的前提依赖性是已知的。将我们的方法应用于50个ProntoQA问题,使用GPT-4o,在检测证明树依赖性方面,我们的方法F1值达到0.806(在谓词确定依赖性方面F1 = 0.885;召回率 = 100%),显著优于自一致性基线(F1 = 0.343;95%的自举置信区间不重叠)。我们还进一步发现,66%正确解决的问题在一致替换下包含至少一个对直接证明树依赖性不敏感的对齐步骤——所有这些都涉及实体引入前提,这是一致替换评估器记录的盲点——这是被动方法无法察觉的“正确答案,错误推理”信号。所有审计证书、原始输出和重现脚本都可在公共GitHub存储库中获取,并且我们讨论了超出形式化、可解析基准的范围限制。

英文摘要

Large language models produce chain-of-thought (CoT) reasoning that appears logically sound yet may not genuinely depend on its stated premises. We introduce interventional grounding audits, a black-box, step-level test of premise dependency: we intervene on a single premise by substituting its target predicate with a fresh symbol, re-run the model, and check whether each reasoning step's normalized conclusion (canonical predicate form) changes. We evaluate on ProntoQA, a synthetic multi-hop deductive reasoning benchmark with gold proof trees, where step-level premise dependencies are known. Applied to 50 ProntoQA problems with GPT-4o, our method achieves F1 = 0.806 on detecting proof-tree dependencies (F1 = 0.885 on predicate-determining dependencies; Recall = 100%), significantly outperforming a self-consistency baseline (F1 = 0.343; 95% bootstrap CIs non-overlapping). We further identify that 66% of correctly-solved problems contain at least one aligned step insensitive to a direct proof-tree dependency under consistent substitution -- all involving entity-introduction premises, a documented blind spot of the consistent-substitution evaluator -- a "right answer, wrong reasoning" signal invisible to passive methods. All audit certificates, raw outputs, and reproduction scripts are available in a public GitHub repository, and we discuss scope limits beyond formal, parsable benchmarks.

CommentsAccepted at the ICLR 2026 Workshop on Logical Reasoning of Large Language Models (https://iclr.cc/virtual/2026/10017466)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑