arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

论用于思维链忠实性的引导向量的泛化性

On the Generalization of Steering Vectors for Chain-of-Thought Faithfulness

Matthew Nguyen, Kyle Cox, Austin Meek, Iván Arcuschin

arXiv 2607.29062首次发表:更新:

发表机构

University of Virginia; University of Delaware; Poseidon Research(弗吉尼亚大学; 特拉华大学; 波塞冬研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对三个模型探究提升思维链忠实性的引导向量的泛化性,发现其效果主要由评估设置决定,且仅对最大型模型有效,引导可减少未被确认的提示使用。

AI 中文摘要

模型能力的提升在很大程度上得益于思维链(Chain of Thought, CoT)的规模化,这对AI安全而言是一个有前景的发展方向——当模型将其推理过程用语言表述出来时,就可以对其进行监控。然而,在某些情况下,模型不会将其推理过程中的重要步骤表述出来。例如,在提示中被暗示错误答案的模型可能会忽略该提示,即便该提示对其结论具有重要作用。当思维链未能披露具有重要作用的推理步骤时,我们将其描述为不忠实。已有研究表明,激活引导(activation steering)是提升思维链忠实性的有效方法。我们拓展了这一研究方向,在提示式问答场景下,针对三个模型(Gemma-3 4B、Qwen-3.5 9B、Gemma-3 12B),研究用于忠实性的引导在提示类型、数据集以及引导向量构建方法之间的泛化效果。尽管引导仅能可靠地提升最大型模型(Gemma-3 12B)的提示确认率,但我们发现,当引导有效时,其效果可广泛泛化至各类提示类型和数据集;在跨提示和跨数据集分析中,效应量主要由评估设置决定,而非向量的训练设置。向量的构建方式影响甚微——四种构建方法(包括一种优化目标未提及特定提示的方法)产生的效应量相近。最后,我们考虑一种可能性:引导会提升提示的显著性,导致更多提示被使用,而非针对表述行为。然而,我们未发现此类证据——引导使提示使用率大致保持不变,同时减少了隐藏提示使用,即未被确认的提示使用。

英文摘要

Model capabilities have improved in large part due to scaling chain of thought. This has been a promising development for AI safety--where models verbalize their reasoning, it is possible to monitor it. However, in some cases, models do not verbalize important steps in their reasoning process. For example, models prompted with a cue suggesting the incorrect answer may fail to acknowledge that cue, even when it appears instrumental to their conclusion. When chain of thought (CoT) fails to disclose instrumental reasoning steps, we describe it as unfaithful. Prior work has shown that activation steering can be a useful method to improve faithfulness in CoT. We extend this line of work by studying how well steering for faithfulness generalizes across cue types, datasets, and methods of constructing the steering vector for three models (Gemma-3 4B, Qwen-3.5 9B, Gemma-3 12B) in a cued question-answering setting. While steering reliably increases cue acknowledgment for only the largest model (Gemma-3 12B), we find that when steering is effective, its effect generalizes broadly across cue types and datasets--in cross-cue and cross-dataset analyses, effect size is determined primarily by the evaluation setting, rather than the vector's train setting. How the vector is built also matters little--four construction methods, including one whose optimization target mentions no specific cue, yield similar effect sizes. Finally, we consider the possibility that steering promotes the salience of the cue and causes greater cue use, rather than targeting verbalization behaviors. However, we find no evidence for this--steering leaves the rate of cue use roughly unchanged while reducing hidden cue use, i.e., cue use that is not acknowledged.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑