发表机构
Northeastern University(东北大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出Localize-and-Detect,一种两阶段黑盒审计方法,通过定位目标任务并搜索偏见内容,有效检测指令微调模型中的任务级投毒,在216个投毒模型上验证了有效性。
AI 中文摘要
指令微调通过在来自一系列任务(如摘要生成和问答)的指令-响应对上训练预训练语言模型,使其适应遵循指令的能力。任务级投毒利用这种任务结构,操纵微调后的模型在特定目标任务上产生攻击者指定的偏见内容,而无需显式的输入触发器。检测此类攻击具有挑战性,因为不存在可识别的显式触发器,目标任务和偏见内容未知,且良性微调本身也会改变模型行为。我们提出定位与检测(Localize-and-Detect),一种用于任务级投毒的两阶段黑盒审计方法,仅需基础模型和微调模型的输出。在第一阶段,我们通过识别微调模型与基础模型在下一次词元分布上差异最大的候选任务来定位目标任务。在第二阶段,我们在候选任务列表中搜索在微调模型响应中反复出现但基础模型响应中不存在的偏见内容。我们在两个模型家族的216个投毒模型上评估了定位与检测,变化了目标任务、投毒模式、投毒预算和偏见内容类型。我们的评估表明,定位与检测能够在多种投毒设置和模型下有效定位目标任务并检测偏见内容,其检测性能与攻击成功率相关,且在干净模型上产生很少的误报。
英文摘要
Instruction fine-tuning adapts a pretrained language model to follow instructions by training it on instruction--response pairs from a collection of tasks, such as summarization and question answering. Task-level poisoning exploits this task structure to manipulate the fine-tuned model into producing attacker-specified biased content on a particular target task, without requiring an explicit input trigger. Detecting such attacks is challenging because there is no explicit trigger to identify, the target task and biased content are unknown, and benign fine-tuning itself changes model behavior. We introduce Localize-and-Detect, a two-stage black-box auditing method for task-level poisoning that requires only outputs from both the base and fine-tuned models. In the first stage, we localize the target task by identifying candidate tasks on which the fine-tuned and base models have the largest differences in their next-token distributions. In the second stage, we search the shortlisted tasks for biased content that repeatedly appears in the fine-tuned model's responses but not in those of the base model. We evaluate Localize-and-Detect on 216 poisoned models across two model families, varying the target task, poisoning mode, poison budget, and type of biased content. Our evaluation demonstrates that Localize-and-Detect can effectively localize target tasks and detect biased content across a range of poisoning settings and models, with detection that tracks attack success and few false positives on clean models.