发表机构
University of Sheffield; Yale University; University of Alabama; Arcadia Impact(谢菲尔德大学; 耶鲁大学; 阿拉巴马大学; 阿卡迪亚影响力)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出ToxScreen基准,研究现实条件下防御者能否恢复LLM的后门触发器,发现令牌查找方法可有效恢复触发器,且易越狱模型是有用信号,相关模型与代码已公开。
AI 中文摘要
随着大语言模型(LLM)被部署到高风险领域,攻击者可能会毒化训练数据以植入后门:即隐藏的触发器,在推理时秘密操纵模型行为。本文研究在现实防御条件下,防御者能否恢复此类触发器,具体条件包括:可白盒访问模型权重、了解目标行为,但无训练数据、无可信参考模型、无触发器知识,且不确定模型是否被毒化。为评估现实场景下防御者恢复触发器的能力,本文发布了ToxScreen基准,包含约800个植入后门的模型,涵盖攻击目标、触发器机制、投毒率、模型规模及后门训练机制。本文断言这些后门质量较高:攻击成功率高、可泛化到未见过的有害输入,且保留干净任务的性能。通过对植入触发器的恢复进行评分,研究发现基于梯度的提示优化无法恢复触发器,而按攻击成功率对候选进行排序的令牌查找方法,可在后门有效的任何场景下恢复触发器。为深入理解这一现象,本文研究了攻击行为与LLM权重的关系,发现后门与越狱采用不同机制策略,防御者可借此过滤越狱。最后,没有方法能可靠地检测出所有后门,但易越狱的模型本身就是异常的,即使未恢复确切触发器,这也是有用信号。本文发布了所有模型及评估代码。
英文摘要
As large language models (LLMs) are deployed in high-stakes domains, adversaries may poison training data to implant backdoors: hidden triggers that covertly manipulate model behavior at inference time. We ask whether a defender can recover such a trigger under realistic affordances, namely white-box access to the weights and knowledge of the behavior of concern, but no training data, no trusted reference model, no knowledge of the trigger, and no certainty that the model is poisoned. To evaluate whether a defender can recover such a trigger under realistic settings, we release ToxScreen, a benchmark of roughly 800 backdoored models spanning attack objectives, trigger mechanisms, poisoning rates, model scales, and backdoor training mechanisms. We also assert that the backdoors are high-quality: they achieve high attack success rates, generalize to unseen harmful inputs, and preserve clean-task performance. Scoring recovery of the planted trigger, we find that gradient-based prompt optimization fails in recovery, whereas a token look-up that ranks candidates by attack-success rate recovers the trigger wherever the backdoor is effective. To understand this more, we study the relationship between attack behaviors and the weights of an LLM. We find a phenomenon whereby backdoors operate via different mechanistic strategies than jailbreaks, allowing defenders to filter jailbreaks. Finally, no method reliably surfaces every backdoor, but a broadly jailbreakable model is itself anomalous, a useful signal even when the exact trigger is not recovered. We release all models and evaluation code