arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过累积熵进行抗后门核心集选择

Anti-Backdoor Coreset Selection via Cumulative Entropy

Qi Zhao, Christian Wressnegger

arXiv 2607.25502首次发表:更新:

AI 中文总结

研究针对神经网络后门的训练阶段防御,将其转化为核心集选择问题,提出用累积熵作选择标准,在各 epoch 去除所选样本,构建良性核心集训练无后门模型,有效减轻后门攻击且对自然性能影响小。

AI 中文摘要

近期针对神经网络后门的训练阶段防御措施是从受污染的训练数据中分离出良性子集,以从中学习无后门模型。本文将此防御策略表述为核心集选择问题,即所谓的“抗后门核心集选择”。由于有毒样本预测不确定性较低且频率低于良性样本,核心集选择自然更关注与良性功能相关的样本。我们使用累积熵作为选择标准来强化这一效果,该指标跟踪训练样本的学习动态,使我们能为核心集选择信息量大的良性样本。此外,我们在每个 epoch 中去除所选样本,以促进良性和有毒样本的可分离性。这产生了一种极其有效的训练阶段防御方法,可构建良性核心集来训练无后门模型。与之前损害自然精度且无法抵御某些攻击的防御方法不同,我们的方法能有效减轻后门攻击,对自然性能影响可忽略不计。

英文摘要

Recent training-time defenses against neural backdoors isolate a benign subset from poisoned training data, to learn a backdoor-free model from it. In this paper, we formulate this defense strategy as a coreset selection problem, giving rise to so-called "Anti-Backdoor Coreset Selection." Since poisonous samples have (a) lower prediction uncertainty and are (b) less frequent than benign samples, coreset selection naturally focuses more on samples associated with benign functionality than the backdoor functionality. We use the Cumulative Entropy as selection criterion to further facilitate this effect. The metric tracks the learning dynamics of training samples and allowing us to select benign samples with high informativeness for the coreset. Additionally, we unlearn the chosen samples in each epoch to facilitate the separability between benign and poisonous samples. Together, this yields an exceptionally effective training-time defense that constructs a benign coreset to train a backdoor-free model. Unlike prior defenses that compromise natural accuracy and fail against certain attacks, our method mitigates backdooring attacks consistently with a negligible impact on natural performance.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑