arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

训练用于自解释忠实性的大语言模型

Training Large Language Models for Self-Explanation Faithfulness

Yeoktatt Cheah, María Pérez-Ortiz, Noah Y. Siegel, Oana-Maria Camburu

arXiv 2607.21090首次发表:更新:

发表机构

Imperial College London(帝国理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出强化学习方法,将忠实性度量改为训练目标,研究模型能否检测及优化影响决策因素以提高自解释忠实性。实验用随机词和用户偏差插入,微调模型在忠实性度量上显著改进,交叉干预泛化有模型依赖等情况,为减少不忠实推理提供路径。

AI 中文摘要

我们提出一种强化学习(RL)方法,直接优化自解释的忠实性,即模型生成的推理准确反映其内部决策过程的程度。现有工作集中于评估忠实性或用推理时提示框架提高大语言模型自解释的可处理性,未提供直接优化模型参数以生成忠实自解释的机制。我们通过将现有忠实性度量修改为RL训练目标来弥合这一差距。研究了模型能否被训练准确检测影响其决策的因素,以及RL能否直接优化这些因素的披露从而提高大语言模型自解释的忠实性。实验采用两种干预类型:随机词插入和用户偏差插入,使用基于Phi-CCT相关度量的逐样本奖励。RL微调的Llama3.1-8B和Qwen3-8B在Phi-CCT忠实性度量上有显著改进,分布内分数从接近零升至高达0.664,在诸如StrategyQA等保留任务上分布外分数达到0.691。交叉干预泛化较弱但更有趣:先验上不会期望仅在随机词插入上训练的模型能泛化到用户偏差短语,但Llama3.1-8B显示了朝此方向的非零迁移。反向方向和Qwen3-8B未复制此结果,表明存在模型依赖和设置依赖效应。最后分析模型行为以排除常困扰RL训练的奖励博弈行为。最终,我们表明模型可被训练隐式识别影响因素并披露它们,为减少大语言模型中不忠实推理提供了一条可扩展路径。

英文摘要

We propose a Reinforcement Learning (RL) method to directly optimize the faithfulness of self-explanations - the extent to which a model's generated reasoning accurately reflects its internal decision-making process. While existing work focuses on evaluating faithfulness or using inference-time prompting frameworks to improve an LLM's self-explanation's tractability, these approaches do not provide a mechanism to directly optimize a model's parameters to generate faithful self-explanations. We bridge this gap by modifying existing faithfulness metrics into an RL training objective. We investigate (1) if models can be trained to accurately detect factors that affect their decisions, and (2) whether RL can directly optimize for the disclosure of these factors thereby improving LLM self-explanations' faithfulness. We experiment with two intervention types: random-word insertions and user-bias insertions, using a per-sample reward derived from the Phi-CCT correlation metric. RL fine-tuned Llama3.1-8B and Qwen3-8B show substantial improvements on the Phi-CCT faithfulness metric, with in-distribution scores rising from near-zero to as high as 0.664, and out-of-distribution scores reaching up to 0.691 on held-out tasks such as StrategyQA. Cross-intervention generalization is weaker but more interesting: a priori we would not expect a model trained only on random word insertions to generalize to user-bias phrases, yet Llama3.1-8B shows non-zero transfer in this direction. The reverse direction and Qwen3-8B do not replicate this, indicating model-dependent and setup-dependent effects we cannot yet explain. Lastly we analyze model behavior to rule out reward gaming behaviors that often plague RL training. Ultimately, we show that models can be trained to implicitly identify influential factors and disclose them, offering a scalable path toward reducing unfaithful reasoning in LLMs.

CommentsTo appear at the ICLR 2026 Workshop on Representational Alignment (Re-Align), 10 pages (long paper)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑