发表机构
Queen Mary University of London; Squirrel AI Learning(伦敦玛丽女王大学; 松鼠AI智适应教育)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多轮LLM,提出答案侧后门攻击:通过良性首轮诱导模型自生成触发器词,使后续有害查询绕过安全拒绝,在4个模型上以5%投毒率达到近100%成功率,并规避输入中心防御。
AI 中文摘要
大语言模型(LLM)的安全对齐仍然容易受到后门攻击的影响。现有的LLM后门几乎都是以输入为中心的:其激活依赖于用户输入中的显式触发器模式,因此现代防护机制旨在净化输入空间。我们通过一种针对多轮对话的新型答案侧后门来挑战这一假设。攻击者不将触发器插入输入,而是使用一个良性的首轮提示自然地诱导模型生成一个特定且看似无害的单词。一旦该单词被合并到对话历史中,这个自生成的单词便成为触发器。当后续出现有害查询时,模型检测到自身的触发器并绕过其安全拒绝机制,而用户输入始终保持完全干净。在四个LLM上,我们的攻击达到了接近完美的攻击成功率,在仅5%的投毒率下接近100%,同时保持了通用实用性和干净输入安全性,并且能够规避主流的以输入为中心的防御措施。表示层面的分析表明,自生成的触发器持续抑制模型的拒绝信号,暴露了当前LLM防御中的一个关键盲点。
英文摘要
Safety alignment in Large Language Models (LLMs) remains vulnerable to backdoor attacks. Existing LLM backdoors are almost all input-centric: activation depends on explicit trigger patterns in the user input, so modern guardrails are built to sanitize the input space. We challenge this assumption with a novel answer-side backdoor for multi-turn dialogue. Instead of inserting the trigger into the input, the adversary uses a benign first-turn prompt to naturally induce the model to generate a specific, seemingly innocuous word. Once merged into the dialogue history, this self-generated word becomes the trigger. When a later harmful query arrives, the model detects its own trigger and bypasses its safety refusal, while the user input stays perfectly clean. Across four LLMs, our attack reaches near-perfect Attack Success Rates, approaching 100\% at only a 5\% poisoning rate, while preserving general utility and clean-input safety, and it evades mainstream input-centric defenses. Representation-level analysis shows that the self-generated trigger consistently suppresses the model's refusal signal, exposing a critical blind spot in current LLM defenses.