arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

高效推理训练并不总是损害思维链的忠实性与可监控性

Efficient Reasoning Training Does Not Always Harm CoT Faithfulness and Monitorability

Samuel Lewis-Lim, Xingwei Tan, Mario Sanger, Zhixue Zhao, Nikolaos Aletras

arXiv 2610.03509首次发表:更新:

发表机构

University of Sheffield; AstraZeneca(谢菲尔德大学; 阿斯利康)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过三种长度压力方法微调模型,发现高效推理训练降低思维链忠实性但可监控性保持稳健,模型仍能承认输入干预的影响。

AI 中文摘要

思维链(CoT)推理使人类能够检查大型语言模型如何得出答案,并监督模型行为。这种推理带来了推理成本的增加,从而推动了高效方法的发展,这些方法训练模型使用更少的令牌来解决问题。然而,一个常见的担忧是,此类训练可能导致模型跳过重要的推理步骤,使得思维链不再忠实地反映模型的决策。目前尚不清楚在实践中这种情况是否发生以及何时发生,因为不同的高效方法以不同的方式对模型的思维链施加长度压力,并且在某些任务上,忠实地解释模型的决策需要比在其他任务上更多的令牌。为了理解这些动态,我们使用三种以不同方式施加长度压力的方法微调了多种模型,即固定生成预算、每个样本的长度目标和组相对长度奖励。我们评估了高效推理如何影响思维链的忠实性(即思维链在相关输入上反映模型决策的程度)以及可监控性(即思维链是否揭示输入干预何时改变输出)。我们发现,它对忠实性和可监控性的影响不同。在大多数设置中,忠实性下降,主要原因是训练后的模型一致性较差。可监控性更为稳健,因为即使思维链明显缩短,模型仍会承认其答案受到的影响。

英文摘要

Chain-of-thought (CoT) reasoning allows humans to inspect how large language models reach their answers, and oversee model behaviour. This reasoning comes at an increased inference cost, motivating efficient methods that train models to solve tasks using fewer tokens. However, a common concern is that such training may cause models to skip important reasoning steps, so the CoT no longer faithfully reflects the model's decision. It is unclear whether or when this occurs in practice, since different efficiency methods apply length pressure to models' CoT in distinct ways, and faithfully explaining a model's decision takes more tokens on some tasks than others. To understand these dynamics, we fine-tune a variety of models with three methods that apply length pressure differently, namely a fixed generation budget, a per-example length target, and a group-relative length reward. We evaluate how efficient reasoning affects CoT faithfulness (i.e., how well the CoT reflects model decisions on related inputs) and monitorability (i.e., whether the CoT reveals when input interventions alter the output). We find that it affects faithfulness and monitorability differently. Faithfulness falls in most settings, primarily because the trained models are less consistent. Monitorability is more robust, as models keep acknowledging the influence on their answer even when the CoT is much shorter.

CommentsUnder Review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑