AI 中文总结
本研究提出新颖探测架构,结合最大欺骗数据集FIBS训练,可高效检测LLM智能体的欺骗与破坏行为,在SHADE-Arena等测试中AUC达98.8%,还能识别内省式欺骗,成果助力前沿监控部署。
AI 中文摘要
近期事件凸显了监控大语言模型(LLM)智能体的挑战以及模型欺骗人类的危险。我们表明,通过探测工具进行的白盒欺骗检测可扩展到前沿监控场景,为此我们收集了迄今为止最大的欺骗数据集用于训练探测工具,并引入了一种新颖的探测架构,该架构可聚合多个层和标记的信息。我们的探测工具在SHADE-Arena中达到了98.8%的AUC,超过了Opus 5.5文本监控基线,且随着底层模型规模扩大,其效能也有所提升。为了将探测工具推向极限,我们在几个仅靠上下文无法确定欺骗的案例中对其进行了测试。在这些我们称为内省式欺骗的案例中,真实情况只能通过仔细引导或深入了解模型的训练数据来确定。在一次此类评估中,我们表明探测工具可区分包含模型真实隐藏目标的文本与其他目标,AUC高达99.7%。我们的探测工具还能轻松检测到知名开放权重模型在政治敏感话题上的欺骗,以及在压力下对自身信念的欺骗。我们发布了名为FIBS的训练数据集,以推动有效探测工具的前沿部署,并鼓励社区用更多欺骗和破坏案例对其进行扩展。
英文摘要
Recent incidents have highlighted the challenge of monitoring LLM agents and the danger of models deceiving people. We show that white-box deception detection via probes can be scaled up to frontier monitoring settings by collecting the largest deception dataset to date for training probes and introducing a novel probe architecture which can aggregate information across many layers and tokens. Our probes achieve 98.8% AUC in SHADE-Arena, surpassing an Opus 5.5 text-monitoring baseline, and show improved efficacy as the underlying model is scaled up. To push our probes to their limit, we test them on several cases where deception cannot be determined from the context alone. In these cases, which we refer to as introspective deception, the ground truth can only be determined through careful elicitation or thorough knowledge of a model's training data. In one such evaluation, we show that probes can distinguish transcripts containing a model's true hidden goal from other goals with an AUC of up to 99.7%. Our probes also readily detect deception on prominent open-weight models which lie about politically sensitive topics, and about their beliefs when put under pressure. We release our training dataset, dubbed FIBS, to help drive frontier deployment of effective probes, and encourage the community to expand upon it with further examples of deception and sabotage.
Comments11 pages main text, 98 pages total; 18 figures, 18 tables. Code and data: https://github.com/AlignmentResearch/caught-in-the-act-probes