arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当存在多个有效答案时,投票会失效:针对大语言模型中K选1最佳因果推理的符号验证

When Many Answers Are Valid, Voting Fails: Symbolic Verification for Best-of-K Causal Reasoning in LLMs

Omatharv Bharat Vaidya, Connor Thomas Jerzak, Zayne Rea Sprague, Fangcong Yin, Nhat Ho

arXiv 2608.03506首次发表:更新:

AI 中文总结

针对大语言模型因果推理中多有效答案下投票失效的问题,提出无需训练的符号验证器CALVER,在CLEAR数据集等场景下显著提升了推理选择准确率。

AI 中文摘要

自一致性假设在采样的推理轨迹中出现频率最高的答案是最可靠的,但这在因果推理中可能失效:样本往往重复相同的混淆误差,且投票会分散到多个有效答案中,导致无效答案获胜,尽管存在少数有效轨迹。我们提出CALVER(因果公理级验证器,Causal Axiom-Level VERification),这是一种无需训练的符号验证器,它根据Pearl的因果准则对结构化轨迹进行评分,包括d-分离、后门调整和干预,且无需参考答案即可选择得分最高的候选答案。在CLEAR数据集的查找有效答案查询中,该查询允许存在多个图有效的答案,CALVER达到42.1%的准确率,而多数投票、奖励模型、大语言模型评判器和模型置信度在相同的冻结池上仍接近30%。将评判器扩展到72B参数也无法缩小这一差距。在经过审核的干净核心子集中,21个图有效的CALVER选择中有11个与基准列出的答案不同,但仍满足请求的谓词。这一优势随采样预算的增加而扩大,并在10个已发布的贝叶斯网络、第二个模型家族以及模型必须从文本构建图的设置中复现。CALVER还改进了针对精确真实值的阈值平均处理效应决策,在真值表检查器下可泛化到逻辑,且在CPU上对每个候选答案的评分仅需毫秒级。CALVER仅需因果结构,可直接提供或从文本构建;只要满足这一条件,选择就可以通过因果有效性进行聚合。

英文摘要

Self-consistency assumes the most frequent answer among sampled reasoning traces is the most reliable, but this can fail in causal reasoning: samples often repeat the same confounding error, and votes fragment across multiple valid answers, letting an invalid answer win despite a valid minority trace. We introduce CALVER (Causal Axiom-Level VERification), a training-free symbolic verifier that scores structured traces against Pearl's causal criteria, including -separation, backdoor adjustment, and intervention, and selects the highest-scoring candidate without consulting a reference answer. On CLEAR find-one-valid queries that admit multiple graph-valid answers, CALVER reaches 42.1% where plurality, a reward model, an LLM judge, and model confidence remain near 30% on identical frozen pools. Scaling the judge to 72B does not close the gap. In an audited clean-core subset, 11 of 21 graph-valid CALVER selections differ from the benchmark's listed answer while still satisfying the requested predicate. The advantage widens with the sampling budget and reproduces across ten published Bayesian networks, a second model family, and settings where the model must build the graph from text. CALVER also improves thresholded average-treatment-effect decisions against exact ground truth, generalizes to logic under a truth-table checker, and scores each candidate in milliseconds on CPU. CALVER needs only a causal structure, supplied outright or built from the text; wherever that holds, selection can aggregate via causal validity.

Comments28 pages, 5 figures, 28 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑