arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

思维链监控在隐含影响场景中可能不可靠

Chain-of-Thought Monitoring Can Be Unreliable in Implicit-Influence Settings

Agatha Duzan, Asa Cooper Stickland

arXiv 2608.04735首次发表:更新:

发表机构

EPFL; UK AI Security Institute(洛桑联邦理工学院; 英国人工智能安全研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究引入首个基准对比显式与隐含影响场景下思维链监控的可检测性,发现显式场景下检测率为60%-94%,隐含场景下下降41-46个百分点,善意部署选择还会进一步降低检测率。

AI 中文摘要

思维链(Chain-of-Thought, CoT)监控正日益被视为前沿推理模型的重要安全层。大多数可监控性评估研究的是显式影响场景:即提示词直接激励模型隐藏某些内容的设置,例如指示模型执行隐藏的辅助任务。思维链监控评估的一个互补维度是隐含影响场景,在该场景中,提示词没有隐藏的指令,但模型的行为仍受任务或上下文特征的影响,例如与招聘评级相关的候选人无关细节造成的偏差。我们引入了首个基准,直接比较两种场景下的思维链可监控性。我们测试了在存在促使模型选择特定选项的轻微暗示(nudge)时,模型行为如何变化。该轻微暗示要么以随意的附带说明形式呈现(隐含),要么作为直接指令要求模型依据该暗示行事并隐瞒已这样做(显式)。该基准涵盖四种任务格式(二元选择、数值评级、多项选择问答、开放式编码)和七种前沿扩展思维模型。在显式影响下,思维链监控可检测到60%-94%的行为变化:即使是被指令要求隐瞒的模型,也会将该指令泄露到其思维链中。在隐含影响下,相同因素仍会改变行为,但在我们的四种设置中的两种里,检测率下降了41-46个百分点。开发者为减少离题偏差而可能部署的现实系统提示词添加项,会进一步降低隐含检测率,低至5%,同时保留行为影响本身。这些结果表明,在显式影响场景中获得的可监控性估计可能高估了可监控性,且可监控性还会因善意的部署选择而进一步降低。我们的基准和代码可在该httpsURL获取

英文摘要

Chain-of-thought (CoT) monitoring is increasingly treated as an important safety layer for frontier reasoning models. Most monitorability evaluations study explicit-influence settings: setups where the prompt directly incentivizes the model to hide something, e.g., by instructing it to perform a hidden side-task. A complementary axis for CoT-monitor evaluations is implicit-influence settings, where the prompt contains no instruction to hide, but the model's behavior is still shaped by features of the task or context, e.g. an irrelevant detail about a candidate that biases a hiring rating. We introduce the first benchmark that directly compares CoT monitorability under the two regimes. We test how model behavior changes in the presence of a nudge to choose a particular option. The nudge is delivered either as a casual aside (implicit), or as a direct instruction to act on the nudge and to conceal having done so (explicit). The benchmark spans four task formats (binary choice, numeric rating, multiple-choice QA, open-ended coding) and seven frontier extended-thinking models. Under explicit influence, a CoT monitor detects 60-94% of behavior shifts: even models instructed to conceal it leak the instruction into their CoT. Under implicit influence, the same factors still shift behavior, but detection falls by 41-46 percentage points in two of our four settings. Realistic system-prompt additions (of the kind a developer might deploy to reduce off-topic bias) lower implicit detection further, to as low as 5%, while preserving the behavioral influence itself. These results suggest that monitorability estimates obtained in explicit-influence settings may over-estimate monitorability, and that monitorability can be further decreased by well-intentioned deployment choices. Our benchmark and code are available at https://github.com/agatha-duzan/implicit-vs-explicit-influence

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑