针对代码任务定制化大语言模型的突破:针对指令后门攻击的自动红队测试
Breaking Customized LLMs for Coding: Automated Red Teaming for Instruction Backdoor Attacks
浏览论文内容
中文总结 AI 辅助
本文提出ARIA自动红队测试框架,可高效生成针对代码定制化LLM的隐蔽带后门指令,攻击成功率达0.945且规避检测能力强,性能优于现有基线攻击。
中文摘要 AI 辅助
大语言模型(LLM)定制平台允许用户通过将指令嵌入系统提示词来构建针对代码智能任务的特定任务模型,无需修改底层模型参数。这些平台降低了定制化LLM的开发门槛,同时也引入了新的攻击面:指令后门攻击,攻击者可通过定制化指令植入隐藏的恶意行为。然而,现有攻击存在两个关键局限:其一,它们通常依赖于平台端或用户端检查可轻易检测到的显式触发模式;其二,它们需要大量人工精力来制作特定任务的带后门指令,限制了可扩展性。本文提出ARIA,一个用于制作针对定制化LLM的隐蔽且有效的带后门指令的自动红队测试框架。ARIA利用攻击者LLM,在目标LLM的结构化反馈引导下,从隐蔽性、干净任务效用、后门有效性三个维度迭代生成并优化带后门指令。我们在三个代码智能任务上,使用四个代表性LLM对ARIA进行评估,并与三个基线攻击进行比较。实验结果表明,ARIA达到了0.945的最高攻击成功率,同时在所有任务中保持最佳的干净任务效用。ARIA还在编程语言间具有良好的泛化能力,且对生成温度保持鲁棒性。此外,ARIA在规避平台端和用户端检测方面显著优于现有攻击,达到最高1.000的假阴性率,并且对现有防御方法仍保持有效,展现出其强大的泛化能力和鲁棒性。
英文摘要
LLM customization platforms allow users to build task-specific models for code intelligence tasks by embedding instructions into system prompts, without modifying the underlying model parameters. While these platforms lower the barrier to developing customized LLMs, they also introduce a new attack surface: instruction backdoor attacks, in which adversaries implant hidden malicious behaviors into customized instructions. However, existing attacks suffer from two key limitations. First, they often rely on explicit trigger patterns readily detected by platform-side or user-side inspection. Second, they require substantial manual effort to craft task-specific backdoored instructions, limiting their scalability. In this paper, we propose ARIA, an automated red-teaming framework for crafting covert and effective backdoored instructions against customized LLMs. ARIA leverages an attacker LLM to iteratively generate and refine backdoored instructions, guided by structured feedback from the target LLM along three dimensions: stealthiness, clean-task utility, and backdoor effectiveness. We evaluate ARIA on three code intelligence tasks, using four representative LLMs, and compare it with three baseline attacks. Experimental results show that ARIA achieves the highest attack success rate of 0.945, while maintaining the best clean-task utility across all tasks. ARIA also generalizes well across programming languages and remains robust to generation temperature. Furthermore, ARIA significantly outperforms existing attacks in evading platform-side and user-side detection, achieving a false negative rate of up to 1.000, and stays effective against existing defense methods, demonstrating its strong generalizability and robustness.