学习隐写术容易,学习隐写推理难
Learning Steganography Is Easy, Learning Steganographic Reasoning Is Hard
浏览论文内容
中文总结 AI 辅助
本研究比较了LLM在强化学习、上下文学习和监督微调下学习隐写推理与相邻能力的难易,发现隐写推理更难,仅在SFT下多数任务可学,但在便利掩护任务下所有方法均可学。
中文摘要 AI 辅助
思维链监控作为一种AI监督与控制的方法,正受到隐写推理可能性的威胁,即大语言模型(LLM)将推理过程隐藏在外观无害的文本中。两个相邻的能力,即隐写消息传递(传递隐藏消息)和编码推理(以难以辨认但未隐藏的格式进行推理),已被证明在真实训练流程中出现的训练压力下(如针对监控器的强化学习)会涌现。这表明隐写推理也可能作为训练的意外副作用而出现。在此,我们比较了模型在三种诱导方法下学习隐写推理及这两种相邻能力的难易程度:强化学习、上下文学习和监督微调(SFT)。对于大多数任务,模型仅在SFT下学习隐写推理,而在所有三种诱导方法下都能学习隐写消息传递和编码推理。即使在SFT下,隐写推理所需的训练量至少是消息传递的两倍,且对于若干模型-任务组合,它根本无法被学习。然而,在一个使得隐藏信息特别方便的掩护任务上,隐写推理可以在所有三种诱导方法下成功学习。因此,隐写推理比隐写消息传递和编码推理难得多,学习后两者并不意味着学习前者。但它仍触手可及:当掩护任务便于隐藏信息时,一个简单的版本可在每种诱导方法下被学习。
英文摘要
Chain-of-thought monitoring as an approach for AI oversight and control is threatened by the possibility of steganographic reasoning, where LLMs conceal their reasoning inside innocuous-looking text. Two neighbouring capabilities, steganographic messaging (passing a concealed message) and encoded reasoning (reasoning in an illegible but unconcealed format), have already been shown to emerge under training pressures that occur in real pipelines, such as reinforcement learning against monitors. This suggests that steganographic reasoning too might arise as an unintended side effect of training. Here, we compare how easily models learn steganographic reasoning and these two neighbouring capabilities across three elicitation methods: reinforcement learning, in-context learning, and supervised fine-tuning (SFT). For most tasks, models learn steganographic reasoning only under SFT, while they learn steganographic messaging and encoded reasoning under all three elicitation methods. Even under SFT, steganographic reasoning requires at least twice as much training as messaging, and for several model-task combinations it is not learned at all. However, on a cover task that makes hiding information especially convenient, steganographic reasoning can be successfully learned under all three elicitation methods. Steganographic reasoning is thus much harder than steganographic messaging and encoded reasoning, and learning the latter two does not imply learning the former. Yet it lies within reach: an easy version is learned under every elicitation method, when the cover task is convenient for hiding information.
发表机构
- Meridian Cambridge
- SAIGE
机构由 AI 辅助整理,请以论文原文为准。