AI 中文总结
本研究提出self-feeding黑盒测试方法,通过将LLM自身输出反馈为输入检测后门,在6个开源权重LLM上实现高模型级检测精度,仅需文本级查询访问即可作为模型安全的低成本初查手段。
AI 中文摘要
任何人都可以将微调后的大语言模型(LLM)上传至公共仓库并声称其是安全的。带有后门的模型在普通输入上表现正常,直到隐藏触发器被触发,而没有训练数据、干净参考权重或触发短语的用户,在使用前没有明确的方法来检查该模型。我们引入并实证评估了self-feeding(自馈)这一黑盒测试方法,该方法将模型自身的输出反馈作为其下一个输入,从而使文本偏离起始提示,向模型微调时使用的数据偏移。我们在6个开源权重LLM(参数规模为3B至15B)上,针对11种攻击类别下的后门,将self-feeding与重复相同提示的基线进行测试,使用20个普通起始提示及最多10步的链。self-feeding在6个模型中的5个中检测到后门,合并精度为92.0%,而相同提示基线仅在120个提示-模型对中的1个上成功;以玩笑请求、算术问题或咖啡配方开头的链均在几步内触发了触发器。每个提示的召回率较低(19.2%),但我们展示了为何在使用多个起始提示后,其在模型层面仍能实现高得多的检测率。我们还报告了该方法的不足:1个模型从未被触发,且self-feeding产生了2个相同提示基线无法产生的误报。将链缩短至4步可在保持所有模型层面检测精度为100%的同时,减少60%的查询量。self-feeding仅需文本级查询访问权限和识别恶意输出的能力,为下载的模型提供了一种低成本的初步检查手段。
英文摘要
Anyone can upload a fine-tuned large language model (LLM) to a public repository and claim it is safe. A backdoored model behaves normally on ordinary inputs until a hidden trigger fires, and a user with no training data, clean reference weights, or the trigger phrase has no clear way to check the model before using it. We introduce and empirically evaluate self-feeding, a black-box test method that feeds a model's own output back as its next input, so the text drifts away from the starting prompt and toward the data the model was fine-tuned on. We test self-feeding against a repeated same-prompt baseline on six open-weight LLMs (3B-15B parameters), each fine-tuned with backdoors spanning eleven attack categories, using twenty ordinary starting prompts and chains of up to ten steps. Self-feeding finds backdoors in five of six models at 92.0\% pooled precision, while the same-prompt baseline succeeds on only one of 120 prompt-model pairs; chains that begin with a joke request, an arithmetic question, or a coffee recipe all reach a trigger within a few steps. Recall per prompt is low (19.2\%), and we show why it still adds up to much higher detection at the model level once several starting prompts are used. We also report where the method falls short: one model was never triggered, and self-feeding produced two false positives that the same-prompt baseline cannot produce. Cutting the chains to four steps keeps every model-level detection at 100\% precision while using 60\% fewer queries. Needing only text-level query access and a way to recognize malicious output, self-feeding offers a cheap first check on a downloaded model.