arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.13760cs.CLcs.AIcs.CVcs.LG

放大并不意味着具有预测性:思考模型中的推理行为

Amplified Does Not Mean Predictive: Reasoning Behaviors in Thinking Models

Jean de Dieu Nyandwi, Leena Mathur, Yonatan Bisk, Robert Hawkins, Graham Neubig

首次发表
浏览论文内容

中文总结 AI 辅助

该研究揭示思考模型存在“放大-提升差距”,即面向推理的训练放大了与正确性关联弱的推理行为,却未放大高提升值的关键行为,推动了过程级推理目标的研究。

中文摘要 AI 辅助

推理模型中哪些推理行为与正确答案相关,面向推理的训练是否会放大这些行为?这种区分十分重要,因为面向推理的训练可能会让推理轨迹看起来更审慎,却不会放大与模型正确性最相关的行为。我们用“行为提升(Behavioral Lift)”来量化这种不匹配,该指标衡量当模型推理轨迹中存在某一行为时,正确性会发生多大变化。在涵盖纯文本和视觉-语言推理的15个模型和6个基准测试中,我们用一套为大语言模型(LLM)和视觉语言模型(VLM)轨迹定义核心行为的分类法,标注了15282条轨迹。我们发现存在“放大-提升差距”:思考模型会大幅放大自我修正、假设检验和不确定性确认,而提升值最高的行为是置信度校准、知识对齐和自我意识。置信度校准是两种模态中正确性的最强积极信号之一,但几乎未被放大;不确定性确认被放大了3至7倍,却与正确性呈弱相关或负相关。我们发现,面向推理的训练不会优先放大提升值最高的行为,这推动了过程级目标的发展,该目标奖励校准和基于事实的推理,而非仅关注表面形式。

英文摘要

Which reasoning behaviors are associated with correct answers in reasoning models, and does reasoning-oriented training amplify those behaviors? This distinction is important because reasoning-oriented training can make traces look more deliberative without amplifying the behaviors most tied to model correctness. We quantify this mismatch with Behavioral Lift, a metric that measures how much correctness changes when a behavior is present versus absent in a model's reasoning trace. Across 15 models and 6 benchmarks spanning text-only and vision-language reasoning, we annotate 15,282 traces with a taxonomy whose core behaviors are defined for both LLM and VLM traces. We find evidence for an Amplification-Lift Gap, in which thinking models strongly amplify self-correction, hypothesis testing, and uncertainty acknowledgment, while the highest-lift behaviors are confidence calibration, knowledge alignment, and self-awareness. Confidence calibration is among the strongest positive signals of correctness in both modalities, yet is barely amplified; uncertainty acknowledgment is amplified by 3--7$\times$, yet is weakly or negatively associated with correctness. We find that reasoning-oriented training does not preferentially amplify the highest-Lift behaviors, motivating process-level objectives that reward calibrated and grounded reasoning rather than surface form alone.

发表机构

  • Carnegie Mellon University(卡内基梅隆大学)
  • Stanford University(斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑