发表机构
University of Washington; Allen Institute for Artificial Intelligence(华盛顿大学; 艾伦人工智能研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究发现预训练数据可通过公共讨论界面被下毒,引入HalfLife方法衡量恶意内容,探索在网络规模下毒预训练语料库的可行性,证明估计毒注入重要性,确立第三方网页内容为攻击语言模型预训练的可能载体。
AI 中文摘要
毒害预训练数据会给语言模型引入难以检测和缓解的有害行为。以往毒害预训练数据的工作大多利用维基百科等既定数据源,未体现预训练语料库的大规模和异质性,且忽视了中毒数据与数据处理管道的交互。本文通过现有网络规模内容注入机制——公共讨论界面,证明了在这种有限设置之外对预训练数据进行中毒攻击是可行的。此外,为衡量网络爬虫和数据处理后是否包含恶意内容,引入了HalfLife,一种用于估计基于网络爬虫的语言模型训练数据中对抗性内容包含情况的新颖分析方法。利用HalfLife探索通过开放讨论界面在网络规模下毒预训练语料库的可行性。分析表明估计预训练数据中是否包含毒注入的重要性,并将第三方网页内容确立为攻击语言模型预训练的可能载体。
英文摘要
Poisoning pretraining data can introduce harmful behaviors to LMs that are difficult to detect and mitigate. Prior work on poisoning pretraining data has largely exploited established data sources such as Wikipedia, which do not represent the large scale and heterogeneity typical of pretraining corpora, and has ignored the interaction between poisoned data and data curation pipelines. We demonstrate that poisoning attacks on pretraining data are feasible beyond this limited setting through an existing web-scale content injection mechanism: public discussion interfaces. Additionally, to measure whether malicious content is included after web crawling and data curation, we introduce HalfLife, a novel analysis for estimating adversarial content inclusion in web-crawl based LM training data. We use HalfLife to explore the feasibility of poisoning pretraining corpora at web scale through open discussion interfaces. Our analysis demonstrates the importance of estimating whether poison injections are included in pretraining data, and establishes third-party webpage content as a possible vector for attacking language model pretraining.