arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.38972cs.CLcs.AI

让大语言模型说出其所想:测量与改进 CoT-可解释性对齐

Making LLMs Say What They Think: Measuring and Improving CoT-Interpretability Alignment

Yihuai Hong, Shauli Ravfogel, Chen Zhao, Eunsol Choi

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出思维链-可解释性对齐(CIA)指标,测量LLM思维链与内部推理的一致性,并通过后训练优化,显著提升忠实度同时保持任务准确性。

中文摘要 AI 辅助

思维链(CoT)轨迹通常被用作大语言模型(LLMs)如何得出其答案的代理。然而,越来越多的证据表明,模型的思维链往往无法反映其内部计算,并且可以在不影响最终答案的情况下被改变。在这项工作中,我们测量并改进LLM思维链中描述推理与其内部计算之间的对齐程度。我们提出了思维链-可解释性对齐(CIA)指标,该指标衡量模型思维链轨迹与其内部推理策略(由可解释性工具检测)之间的一致性。我们在三个任务(两跳问答、提示干预和整数乘法)上评估CIA,涉及三个LLM,发现LLM在所有任务中表现出有限的对齐(44.8-75.9%)。随后,我们尝试通过后训练来改进CIA,将任务准确性和参数化忠实度信号同时作为奖励。实验表明,我们可以在保持或提高任务准确性的同时,显著提高思维链的参数化忠实度。我们提供了丰富的分析,例如其泛化模式。我们的工作既提供了一个审计思维链参数化忠实度的框架,也提供了一条使模型显式推理更值得信赖的途径。代码和数据可在以下网址获取:此 https URL。

英文摘要

Chain-of-thought (CoT) traces often serve as a proxy for how Large Language Models (LLMs) arrive at their answers. However, growing evidence shows that models' CoT often fails to reflect their internal computations and can be changed without affecting their final answers. In this work, we measure and improve the alignment between the reasoning described in an LLM's CoT and what it computes internally. We propose CoT-Interpretability Alignment (CIA), a metric that measures the agreement between a model's CoT traces and its internal reasoning strategies as detected by interpretability tools. We evaluate CIA on three tasks (two-hop question answering, hint intervention, and integer multiplication) across three LLMs, finding that LLMs exhibit limited alignment across all tasks (44.8-75.9%). We then experiment with improving CIA via post-training, setting both the task accuracy and parametric faithfulness signals as a reward. Experiments show that we can substantially improve CoT parametric faithfulness while maintaining or improving the task accuracy. We provide rich analysis, such as their generalization patterns. Our work provides both a framework for auditing CoT parametric faithfulness and a pathway toward making models' explicit reasoning more trustworthy. Code and data are available at https://github.com/yihuaihong/CIA-minimal-repro.

发表机构

  • New York University(纽约大学)
  • NYU Shanghai(上海纽约大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑