arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32496cs.CL

局部可靠,全局不足:多跳推理中的局部-全局差距

Locally Sound, Globally Insufficient: The Local-Global Gap in Multi-Hop Reasoning

Bohao Chu, Hendrik Damm, Qianli Wang, Hui Wang, Shuning Zhang, Christoph M. Friedrich, Norbert Fuhr

首次发表
浏览论文内容

中文总结 AI 辅助

针对多跳推理中局部可靠但全局不充分的问题,提出E-Closure方法,在训练中监督证据支持、问题对齐和答案闭合三类依赖,显著提升准确率与轨迹可靠性并降低LGG率。

中文摘要 AI 辅助

可靠的多跳推理需要的不仅仅是局部支持的步骤:一个推理轨迹可以在每一步推理中都是可靠的,但仍然无法整体回答该问题。我们将这种失败情形称为局部-全局差距(LGG),即轨迹在局部是可靠的,但在全局上是不充分的。局部可靠性要求每一步都得到可用证据和先前步骤的支持,而全局充分性则要求推理轨迹与问题对齐并确立所提交的答案。在对三个多跳问答基准和三个模型的2,598个回答进行人工裁决的诊断中,我们发现LGG出现在每一个基准-模型组合中,并占整体全局不充分回答的近一半。然而,传统的基于证据检查声明的忠实性验证器在很大程度上会遗漏这些失败:在保留至少95%可靠轨迹的阈值下,对LGG案例的召回率显著低于对局部不可靠轨迹的召回率。为了解决这些失败,我们形式化了可靠推理的三个依赖关系:证据到步骤的支持、问题到轨迹的对齐以及轨迹到答案的闭合。我们不是事后诊断,而是引入E-Closure在训练期间监督这些依赖关系,将对受支持的原回答和反事实回答的生成监督与双向切换约束相结合。在三个基准和三个骨干网络上取平均,现有的微调基线在基础模型上提高了准确性和局部可靠性,但以牺牲全局充分性为代价。E-Closure两者都提高了:在所有微调方法中,它实现了最高的平均准确性(92.8%)和轨迹可靠性(89.0%),同时产生了最低的LGG率(6.2%)。

英文摘要

Reliable multi-hop reasoning requires more than locally supported steps: a trace can be sound at every reasoning step yet still fail to answer the question as a whole. We call this failure regime the local-global gap (LGG), in which the trace is locally sound yet globally insufficient. Local soundness requires each step to be supported by the available evidence and preceding steps, whereas global sufficiency requires the reasoning trace to align with the question and establish the submitted answer. In a human-adjudicated diagnostic of 2,598 responses across three multi-hop QA benchmarks and three models, we find that the LGG occurs in every benchmark-model combination and accounts for nearly half of globally insufficient responses overall. However, conventional faithfulness verifiers that check claims against the evidence largely miss these failures: at thresholds retaining at least 95% of reliable traces, recall for LGG cases is substantially lower than that for locally unsound traces. To address these failures, we formalize three dependencies for reliable reasoning: evidence-to-step support, question-to-trace alignment, and trace-to-answer closure. Instead of post-hoc diagnosis, we introduce E-Closure to supervise these dependencies during training, combining generation supervision on supported original and counterfactual responses with bidirectional switching constraints. Averaged over three benchmarks and three backbones, existing fine-tuning baselines improve accuracy and local soundness over the base models, but at the cost of global sufficiency. E-Closure improves both: among all fine-tuned methods, it achieves the highest average accuracy (92.8%) and trace reliability (89.0%) while yielding the lowest LGG rate (6.2%).

补充信息

↑