思维链可监控性中“看不见即想不到”吗?
Does Out-of-Sight Equal Out-of-Mind in CoT Monitorability?
浏览论文内容
中文总结 AI 辅助
本研究对比显式与潜在CoT的可监控性,发现其更多取决于任务属性和模型内部访问程度,而非推理模式。
中文摘要 AI 辅助
思维链(Chain-of-thought, CoT)推理为大型语言模型(Large Language Models, LLMs)的决策提供了窗口,通过读取推理轨迹可监控目标行为,推动了CoT可监控性相关研究。然而,潜在CoT方法用少量连续状态替代显式token,降低了推理成本,但移除了监控所依赖的可读轨迹。监控需通过替代方式访问模型,如探测其激活或把潜在状态转译为文本,但这些替代方式能保留多少可监控性尚不明确。本研究以基于提示的干预设置为研究对象,该设置是模型利用偏差输入线索(如意外泄露的答案或用户陈述的信念)却不承认这些线索的行为的代理,将提示依赖性作为可监控性目标,在数学推理和问答任务中对比不同推理模式下的监控器,包括显式CoT、弱监督潜在CoT和强监督潜在CoT。研究发现,在该设置下,可监控性更多取决于任务属性(如正确答案是否约束支撑推理)和对模型内部的访问程度,而非推理模式。
英文摘要
Chain-of-thought (CoT) reasoning offers a window into the decision-making of large language models (LLMs), which can be monitored for target behaviors by reading the reasoning trace, motivating work on CoT monitorability. Latent CoT approaches, however, replace the explicit tokens with a small number of continuous states, lowering inference costs but removing the readable trace this monitoring relies on. Monitoring then requires alternative access to the model, such as probing its activations or verbalizing the latent states back into text, but how much monitorability these alternatives preserve is unclear. We study this question with a hint-based intervention setup, a proxy for behaviors where models exploit biasing input cues, e.g., an inadvertently leaked answer or a belief stated by the user, without acknowledging them. Taking hint-reliance as the monitorability target, we compare monitors across reasoning modes, from explicit CoT to weakly- and strongly-supervised latent CoT, on math reasoning and question answering. We find that, in this setup, monitorability depends more on properties of the task (such as whether the correct answer constrains the supporting reasoning) and the level of access to model internals than on the reasoning mode.
发表机构
- University of Amsterdam(阿姆斯特丹大学)
- University of Edinburgh(爱丁堡大学)
机构由 AI 辅助整理,请以论文原文为准。