arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

曲线推理 II:休眠智能体几何——将可解释性扩展至探测技术之外

Curved Inference II: Sleeper Agent Geometry - Extending Interpretability Beyond Probes

Rob Manson

arXiv 2608.24037首次发表:更新:

AI 中文总结

本文扩展了Anthropic的休眠智能体研究,提出曲线推理框架,引入语义表面积度量,通过多轮上下文窗口分析自然欺骗推理的几何特征,为线性检测失效时提供了可扩展的无监督检测路径。

AI 中文摘要

本文扩展了Anthropic的休眠智能体研究[1],该研究表明人工后门会在安全训练中持续存在,且可通过线性探测器以99%以上的准确率检测到[2]。但基于探测器的检测依赖于线性可分性,这可能是后门插入的人为产物,而非自然出现的欺骗性对齐的属性。通过自然训练产生的复杂欺骗行为不太可能产生此类便捷的线性信号。我们引入了一种使用多轮上下文窗口的自然主义方法,该方法模拟了现实的欺骗推理,无需人工触发或有监督的后门插入。我们不采用二元触发-响应模式,而是研究语义复杂性如何通过上下文的逐步发展而出现。在我们的曲线推理(Curved Inference)框架基础上,我们分析了曲率、显著性,并引入了语义表面积(A'),这是一种新的表征工作度量,用于捕获未归一化残差空间中意义构建的幅度和方向变化。在无后门、无标签、无探测器的情况下,我们将该框架应用于自然主义欺骗提示,并通过大语言模型(LLM)共识对模型输出进行分类。几何结构可可靠预测语义分类,五种提示策略和两个模型家族的表面积存在统计学显著差异。关键的是,测量精度可揭示被分类噪声掩盖的几何特征——部分策略从无显著性(p=0.555)提升至显著(p=0.048)。这验证了复杂推理会产生内在几何模式,即便检测似乎失败时仍会持续存在,表明推理本身的形状编码了语义模式,无论模型是否学会抑制欺骗的线性指标,这为线性方法失效时提供了一种可扩展的无监督检测路径。

英文摘要

This paper extends Anthropic's Sleeper Agents research [1], which showed artificial backdoors persist through safety training & can be detected by linear probes with >99% accuracy [2]. However, probe-based detection relies on linear separability that may be an artefact of backdoor insertion rather than a property of naturally occurring deceptive alignment. Sophisticated deceptive behaviours emerging through natural training are unlikely to produce such convenient linear signals. We introduce a naturalistic methodology using multi-turn context windows that simulates realistic deceptive reasoning without artificial triggers or supervised backdoor insertion. Rather than binary trigger-response patterns, we examine how semantic complexity emerges through gradual context development. Building on our Curved Inference framework, we analyse curvature, salience, & introduce semantic surface area (A'), a new metric of representational work capturing both the magnitude & directional change of meaning construction in unnormalised residual space. Without backdoors, labels, or probes, we apply this framework to naturalistic deceptive prompts & classify model outputs via LLM consensus. Geometric structure reliably predicts semantic classification, with statistically significant differences in surface area across five prompt strategies & two model families. Critically, measurement precision can reveal geometric signatures hidden by classification noise - some strategies improve from non-significant (p = 0.555) to significant (p = 0.048). This validates that sophisticated reasoning creates intrinsic geometric patterns that persist even when detection appears to fail, suggesting the shape of inference itself encodes semantic patterns regardless of whether models have learned to suppress linear indicators of deception - a scalable, unsupervised path for detection when linear methods fail.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑