发表机构
Bonn-Aachen International Center for Information Technology, University of Bonn; Lamarr Institute for Machine Learning and Artificial Intelligence(波恩大学波恩-亚琛国际信息技术中心; 拉马尔机器学习与人工智能研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出无监督检测方法激活匹配微调,无需预先了解触发器或目标行为,可在不牺牲隐藏行为的情况下,可靠检测大语言模型中的后门等隐藏行为。
AI 中文摘要
大语言模型会隐藏仅在特定狭窄条件下激活的隐藏行为,例如后门触发器、休眠智能体部署线索、沙袋行为(sandbagging)或主题条件审查。若事先不知道要查找的内容,这类行为很难被检测到。我们提出激活匹配微调(activation-matched finetuning),这是一种无监督检测方法,无需预先了解触发器或目标行为。给定一个可疑模型和一个公开可用的锚定模型,我们对锚定模型进行微调,使其在小型良性语料库上复现可疑模型的激活值,并通过两个模型之间的残差对每个评估提示进行评分。由于没有任何良性语料库能覆盖稀疏的触发器区域,参考模型学习到的是良性计算,而非隐藏行为。因此,触发器提示——以及关键的语义邻居——会产生较大的残差,向防御者表明存在异常行为。我们在第三方模型和自定义模型上测试该方法,结果显示激活匹配微调可可靠地揭示隐藏行为。此外,我们通过实验研究了一种天然的防御感知攻击,发现该攻击在不牺牲行为本身的情况下,无法抑制我们的检测方法。
英文摘要
Large language models can hide hidden behaviors that activate only under narrow conditions, such as backdoor triggers, sleeper-agent deployment cues, sandbagging, or topic-conditioned censorship. Such behaviors are difficult to detect without prior knowledge what to look for. We present activation-matched finetuning, an unsupervised detection method that assumes no knowledge of the trigger or the target behavior. Given a suspect model and a publicly available anchor, we finetune the anchor to reproduce the suspect's activations on a small benign corpus, and score each evaluation prompt by the residual between the two models. Since no benign corpus covers the sparse trigger region, the reference learns the benign computation but not the hidden behavior. Therefore, trigger prompts -- and, crucially, their semantic neighbors -- incur a large residual that signal the presence of unusual behavior to the defender. Testing our method across third-party models and custom models, activation-matched finetuning surfaces hidden behavior reliably. Furthermore, we empirically consider a natural defense-aware attack and showcase that it fails to suppress our detection method without sacrificing the behavior itself.
Commentsunder review