arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

你看到的错误并非你犯下的错误:面向推理错误定位的进展感知推理起源

The Error You See Is Not the Error You Made: Progression-aware Reasoning Origin for Reasoning Error Localization

Yiguo Wang, Ziyuan Yang, Yi Zou, Dan Lin, Rongsheng Li, Yi Zhang

arXiv 2609.33297首次发表:更新:

AI 中文总结

针对多步LLM推理验证中错误定位不准的问题,提出无训练框架PRO,通过联合建模上下文支持与后续兼容性并利用干预证据区分错误起源,在多项任务上超越强基线。

AI 中文摘要

验证多步LLM推理需要的不仅仅是判断一条推理轨迹是否正确:一个有用的验证器应当识别推理首次出错的位置。然而,现有的整体性方法几乎不提供位置证据,而前向顺序验证往往将第一个被拒绝的步骤视为错误来源。在错误传播的情况下,这一假设可能失效,因为较早的错误可能在局部看似合理,只有通过其下游后果才变得可观察。因此,我们将推理验证重新构想为一个进展感知的错误来源定位问题:我们不仅询问推理轨迹在何处首次显得不一致,还询问哪个更早的步骤最能解释这种不一致是如何沿轨迹产生的。基于这一观点,我们提出了进展感知推理起源(Progression-aware Reasoning Origin, PRO),一个用于首错定位的无训练框架。PRO联合建模来自前文的上游支持与对后续推理的下游兼容性,选择性地精炼这些信号不一致的区域,并最终执行基于检测器条件的源归因,利用基于干预的证据来区分真正的错误起源与其传播的表现形式。我们进一步形式化了前向拒绝与结构暴露之间的差距,表明仅凭上游侧证据在错误传播下不足以实现可靠的定位。在开放形式、医学和结构化推理任务上的实验表明,与强验证基线相比取得了一致的改进,支持进展感知源归因作为推理验证的一种更忠实的表述。

英文摘要

Verifying multi-step LLM reasoning requires more than determining whether a trace is correct: a useful verifier should identify where the reasoning first goes wrong. However, existing holistic methods provide little positional evidence, while forward sequential verification often treats the first rejected step as the error source. Under error propagation, this assumption can fail, since an earlier mistake may remain locally plausible and become observable only through its downstream consequences. We therefore rethink reasoning verification as a progression-aware error-source localization problem: rather than asking only where a reasoning trace first appears inconsistent, we ask which earlier step best explains how that inconsistency emerges along the trajectory. Based on this view, we propose Progression-aware Reasoning Origin (PRO), a training-free framework for first-error localization. PRO jointly models incoming support from the preceding context and outgoing compatibility with subsequent reasoning, selectively refines regions where these signals disagree, and finally performs detector-conditioned source attribution with intervention-based evidence to distinguish the true error origin from its propagated manifestations. We further formalize the gap between forward rejection and structural exposure, showing why incoming-side evidence alone is insufficient for reliable localization under error propagation. Experiments across open-form, medical, and structured reasoning tasks demonstrate consistent improvements over strong verification baselines, supporting progression-aware source attribution as a more faithful formulation of reasoning verification.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑