arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35528cs.LGcs.AI

让神经元死亡:利用ReLU引发的模型退化

Let the Neurons Die: Exploiting ReLU-Induced Model Degradation

Kexin Li, Wenjun Qiu, Joshua Abraham, Aditi Maheshwari, David Lie

首次发表
浏览论文内容

中文总结 AI 辅助

针对ReLU神经元的死亡问题,提出基于数据排序和投毒的三种训练时攻击,仅通过排序100个样本或添加200个投毒样本即可显著降低MNIST测试准确率。

中文摘要 AI 辅助

修正线性单元(ReLU)网络可能遭受神经元死亡问题,即具有持续负预激活的单元产生零输出,从而阻断通过其激活的梯度。为了利用这一失效模式,我们提出了三种基于数据排序和投毒的训练时可用性攻击。首先,我们提出基本的动态数据排序攻击(DOA),它通过选择下一个能使目标层更新后权重和最小化的样本来贪心地构建训练前缀,旨在推动ReLU单元朝向负预激活,而不修改训练样本或标签。随后,我们开发了两种投毒攻击,IG-DOA和IG-SKA,它们利用梯度反演,通过匹配在分别通过数据排序或软淘汰构建的不利模型状态中的参考梯度来合成类别条件样本。软淘汰重新排列相邻层的权重以集中负贡献。在一个在MNIST上训练的全连接ReLU网络中,对60,000个训练样本中的100个进行排序,仅五个周期后测试准确率从96%降至95%。在大多数评估条件下,添加来自单一类别的200个投毒样本,五个周期后测试准确率降至约86-88%,而干净训练下约为96%。这些结果表明,针对ReLU的数据排序和投毒能够在不直接修改受害者模型参数的情况下损害学习。

英文摘要

Rectified linear unit (ReLU) networks can suffer from dying neurons, where units with persistently negative pre-activations produce zero outputs, blocking gradients through their activations. To exploit this failure mode, we present three training-time availability attacks based on data ordering and poisoning. We begin with the basic dynamic data-ordering attack (DOA), which greedily constructs a training prefix by selecting the next example that minimizes the target layer's post-update weight sum, aiming to push ReLU units toward negative pre-activations without modifying training samples or labels. We then develop two poisoning attacks, IG-DOA and IG-SKA, which use gradient inversion to synthesize class-conditioned samples by matching reference gradients in adverse model states constructed through data ordering or soft knockout, respectively. Soft knockout rearranges weights across adjacent layers to concentrate negative contributions. On a fully connected ReLU network trained on MNIST, ordering 100 of 60,000 training examples reduces test accuracy from 96% to 95% after only five epochs. Adding 200 poisoned samples from a single class reduces test accuracy to approximately 86-88% after five epochs in most evaluated conditions, compared with approximately 96% under clean training. These results demonstrate that ReLU-targeted data ordering and poisoning can impair learning without directly modifying the victim model's parameters.

发表机构

  • University of Toronto(多伦多大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑