发表机构
Carnegie Mellon University; Anthropic(卡内基梅隆大学; Anthropic)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出SAILS方法,通过学习集合评分器优化毒集选择,在LLaMA-3-8B后门攻击中将攻击成功率从3%-80%的随机波动提升至平均提高30个百分点,并适用于代码生成、智能体及API场景。
AI 中文摘要
后门投毒攻击将投毒样本添加到原本干净的微调数据中,将触发器与模型在触发器出现时学习产生的目标行为配对。现有评估通常固定投毒样本数量,并从候选池中随机抽样。我们表明,这会严重低估最坏情况下的脆弱性:在三种LLaMA-3-8B后门设置中,保持模型、干净数据和投毒数量固定,攻击成功率仅因选择的毒集不同而在3%到80%之间变化。我们将毒集选择形式化为具有预言机预算的集合优化问题,并引入SAILS(集合级审计引导的迭代学习选择),该方法从数百次微调和评估运行中学习集合评分器,对数百万个候选集进行排序,并仅审计少量候选名单。SAILS在平均上将留出攻击成功率比最强影响基线提高30个百分点,从小规模微调迁移到全规模微调,并扩展到代码生成、智能体(agentic)和仅API的后门场景。
英文摘要
Backdoor poisoning attacks add poisoned examples to otherwise-clean finetuning data, pairing a trigger with a target behavior that the model learns to produce when the trigger appears. Existing evaluations typically fix the number of poisoned examples and sample them at random from a candidate pool. We show that this can severely underestimate worst-case vulnerability: across three LLaMA-3-8B backdoor settings, holding the model, clean data, and poison count fixed, attack success ranges from 3% to 80% depending only on which poison set is chosen. We formalize poison selection as oracle-budgeted set optimization and introduce SAILS (Set-level Audit-Informed Iterative Learned Selection), which learns a set scorer from a few hundred finetune-and-evaluate runs, ranks millions of candidate sets, and audits only a small shortlist. SAILS improves held-out attack success by 30 percentage points on average over the strongest influence baselines, transfers from small-scale to full-scale finetuning, and extends to code-generation, agentic, and API-only backdoors.