arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35074cs.CL

当信心过早上升:通过过早答案承诺检测捷径推理

When Confidence Rises Too Early: Detecting Shortcut Reasoning via Premature Answer Commitment

  • Queen Mary University of London(伦敦玛丽女王大学)
  • Tongyi Lab, Alibaba Group(阿里巴巴集团通义实验室)

机构由 AI 辅助整理,请以论文原文为准。

Zhaohan Zhang, Junjie Liu, Chengzhengxu Li, Chen Shen, Xiaoming Liu, Chao Shen, Jieping Ye, Ziquan Liu, Ioannis Patras

中文总结 AI 辅助

提出ConfLens框架与DACS分数,通过追踪推理中答案信心过早上升来检测LLM的捷径推理,在数学和代码任务上F1提升超4.3%,并改善奖励模型偏好。

中文摘要 AI 辅助

大型语言模型(LLM)的推理轨迹通常被视为其内部推理的言语化描述。然而,此类轨迹可能不忠实:模型可能依赖捷径来得出答案,然后通过看似连贯的思维链对决策进行事后合理化。检测这种捷径推理具有挑战性,因为现有的监控器和验证器主要检查文本痕迹或最终结果,而非模型在生成过程中对其答案信念的发展方式。我们引入了ConfLens,一个在推理过程中追踪最终答案信心演变的框架。在三种捷径推理设置中,我们观察到一种常见的过早信心模式,即捷径样本在早期推理阶段就对最终答案表现出高度信心。然而,现有的信心估计方法在检测这种行为时表现出有限的泛化性、可靠性或效率。因此,我们提出了分布性答案承诺分数(DACS),一种分布性信心估计器,用于衡量模型在每个推理步骤中对答案承诺的概率分布的熵。DACS捕捉模型答案信念的集中程度,无需真实答案或任务特定验证器。我们进一步将ConfLens的检测结果转换为可解释的信号,用于奖励模型以减少其对捷径推理的偏好。在数学和代码推理任务上的实验表明,与强基线相比,使用DACS的ConfLens将捷径推理检测的F1分数提高了超过4.3%,并减少了奖励模型偏好中忠实性与正确性之间的不匹配。

英文摘要

The reasoning trajectory of a Large Language Model (LLM) is often treated as a verbalized description of its internal reasoning. However, such trajectories can be unfaithful: a model may rely on shortcuts to reach an answer and then post-rationalize the decision with a seemingly coherent chain of thought. Detecting this shortcut reasoning is challenging because existing monitors and verifiers mainly inspect textual traces or final outcomes, rather than how the model's belief in its answer develops during generation. We introduce ConfLens, a framework that tracks the evolution of confidence in the final answer throughout reasoning. Across three shortcut reasoning settings, we observe a common pattern of premature confidence, where shortcut samples become highly confident in the final answer at early reasoning stages. Existing confidence estimation methods, however, show limited generalizability, reliability, or efficiency for detecting this behavior. We therefore propose the Distributional Answer Commitment Score (DACS), a distributional confidence estimator that measures the entropy of the model's probability distribution over answer commitment at each reasoning step. DACS captures how concentrated the model's answer belief is without requiring ground-truth answers or task-specific verifiers. We further convert ConfLens detection results into interpretable signals for reward models to reduce their preference for shortcut reasoning. Experiments on mathematical and code reasoning tasks show that ConfLens with DACS improves shortcut reasoning detection by over 4.3% F1 compared with strong baselines and reduces the mismatch between faithfulness and correctness in reward model preferences.

补充信息

↑