SAGE:基于相似性的从已验证示例中清洗投毒训练数据
SAGE: Similarity-Based Cleaning of Poisoned Training Data from Verified Examples
浏览论文内容
中文总结 AI 辅助
针对数据投毒攻击,提出SAGE方法,利用少量已验证示例(含干净与投毒)通过相似性加权预测检测投毒,在七种攻击基准上验证了有效性。
中文摘要 AI 辅助
随着机器学习越来越依赖公开的、不可信的数据源,数据投毒攻击(向训练数据中注入恶意示例以诱导对选定目标的错误分类)构成了日益增长的威胁。现有的防御方法要么假设关于哪些示例被投毒没有任何真实标签信息,要么假设可以访问大量已验证为干净的示例。满足后一种假设会带来显著成本,因为可靠的验证可能非常耗费资源或人力。对于干净标签攻击,这种成本尤其高,因为投毒示例在视觉上与干净数据无法区分。由于要求大量已验证示例不切实际,我们提出依赖一小部分已验证示例,包括干净和投毒的示例,即每个示例都通过取证专家的检查验证为干净或投毒。挑战在于基于如此小的已验证示例集检测投毒,以至于大多数分类模型会过拟合。为应对这一挑战,我们提出了基于相似性的真实标签驱动排除方法(SAGE),该方法在单独的数据集上训练通用特征提取器,然后基于已验证集使用非参数、相似性加权的预测来标记投毒的训练示例。在针对七种干净标签攻击方法的标准基准上,我们证明即使访问少量已验证的投毒示例也能提供显著优势。我们还发现,已验证干净示例在各类别中的分布比已验证示例的数量更重要。
英文摘要
As machine learning increasingly relies on public, untrusted data sources, data poisoning attacks, which inject malicious examples into training data to induce misclassification of a chosen target, pose a growing threat. Existing defenses either assume zero ground-truth information about which examples are poisoned, or they assume access to a large set of examples verified to be clean. Satisfying the latter assumption incurs significant cost since reliable verification can be very resource- or labor-intensive. This cost is particularly high for clean-label attacks, where poisoned examples are visually indistinguishable from clean data. Since requiring a large set of verified examples is impractical, we propose relying on a small set of verified examples including both clean and poisoned ones, i.e., each example verified either to be clean or poisoned through inspection by a forensic expert. The challenge is then to detect poisons based on a set of verified examples that is so small that most classification models would overfit. To address this challenge, we propose Similarity-based Approach for Ground-truth-driven Exclusion (SAGE), which trains a generic feature extractor on a separate dataset and then flags poisoned training examples using a non-parametric, similarity-weighted prediction based on the verified set. On standard benchmarks against seven clean-label attack methods, we demonstrate that having access to even a handful of verified poisoned examples provides a substantial advantage. We also find that the distribution of verified clean examples across classes matters more than the number of verified examples.
发表机构
- Pennsylvania State University(宾夕法尼亚州立大学)
- University of Louisville(路易斯维尔大学)
- Washington University in St. Louis(圣路易斯华盛顿大学)
机构由 AI 辅助整理,请以论文原文为准。