arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

推理模型中隐藏指令的选择性披露:行为不对称性与引导

Selective Disclosure of Hidden Directives in Reasoning Models: Behavioral Asymmetry and Steering

Zimo Shi, Xander Tifft, Wen Xing

arXiv 2608.29070首次发表:更新:

发表机构

SPAR Research; Yale University(SPAR研究院; 耶鲁大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对8个前沿推理模型,发现其CoT对恶意隐藏指令的披露概率高于良性指令,存在行为不对称性,且共享隐藏方向的差异化激活是该不对称性的来源。

AI 中文摘要

思维链(Chain-of-thought,CoT)推理轨迹正日益被提议作为AI监督的一种机制:从原则上讲,检查模型推理过程的监控器能够检测到仅从输出中无法察觉的不当行为。这一假设的前提是,无论给定何种指令,CoT都能展现出模型被指示执行的任务。我们从两个维度验证了这一假设。首先,我们引入了指令合规性差距(Instruction-Compliance Gap,ICG):当隐藏系统提示指令为恶意时,模型的CoT明确引用该指令的概率,与该指令为良性时的概率之差。在100个任务对和来自5个系列的8个前沿推理模型上,我们发现了一致的不对称披露现象:Qwen3-14B(Wilcoxon检验p=0.0001,+13.9个百分点)、Qwen3-32B(p=0.0011,+13.0个百分点)、Qwen3-235B(p=0.035,+5.8个百分点)存在更高的恶意隐藏指令泄露概率,MiniMax-M2.5和DeepSeek-R1也有类似结果。该检测器在两次独立盲标注过程中精度达100%,仅读取推理轨迹的LLM监控器针对无指令对照组,在所有8个模型中都重现了这种不对称性,且在检测器归类为干净的恶意轨迹中,有82%能识别出具体指令。其次,通过对比激活添加(Contrastive Activation Addition)从MiniMax-M2.5中提取的引导向量,可因果性地诱导对裸提示的隐藏,并抑制原本会产生的隐藏,在预注册设计下于Qwen3-14B中也得到了复现。良性和恶意衍生的隐藏向量高度相似(MiniMax-M2.5中余弦相似度为0.804,Qwen3-14B中为0.970),这意味着在这些模型中,披露不对称性源于共享隐藏方向的差异化激活,而非独立机制。

英文摘要

Chain-of-thought (CoT) reasoning traces are increasingly proposed as a mechanism for AI oversight: a monitor inspecting a model's reasoning can, in principle, detect misbehavior invisible from outputs alone. This assumes CoT surfaces what a model is instructed to do regardless of the instructions given. We test this assumption along two axes. First, we introduce the Instruction-Compliance Gap (ICG): the difference in probability that a model's CoT explicitly references a hidden system prompt directive when that directive is malign versus benign. Across 100 task pairs and 8 frontier reasoning models from 5 families, we find consistent asymmetric disclosure, a higher probability of leaking malign hidden instructions than benign ones, in Qwen3-14B (Wilcoxon $p=0.0001$, $+13.9$pp), Qwen3-32B ($p=0.0011$, $+13.0$pp), Qwen3-235B ($p=0.035$, $+5.8$pp), and similar results with MiniMax-M2.5 and DeepSeek-R1. The detector has 100% precision against two independent blinded labelling passes, and an LLM monitor reading only the reasoning trace reproduces the asymmetry in all 8 models against directive-free controls, identifying the specific directive in 82% of malign traces which the detector classifies as clean. Second, steering vectors extracted in MiniMax-M2.5 via Contrastive Activation Addition causally induce hiding from bare prompts and suppress it from prompts that would otherwise produce it, replicating in Qwen3-14B under a pre-registered design. Benign and malign-derived hiding vectors are highly similar (cosine $0.804$ in MiniMax-M2.5; $0.970$ in Qwen3-14B), implying that in these models the disclosure asymmetry arises from differential activation of a shared hiding direction rather than separate mechanisms.

Comments22 pages, including appendices

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑