RL-FAT:用于公平对抗训练的强化学习
RL-FAT: Reinforcement Learning for Fair Adversarial Training
- University of Mannheim(曼海姆大学)
- MPI for Informatics(马克斯·普朗克信息学研究所)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对对抗训练存在的类别鲁棒性失衡问题,提出受强化学习启发的RL-FAT框架,结合强化驱动自适应与公平强调正则化,在提升鲁棒性的同时大幅降低类别鲁棒性失衡。
AI中文摘要:
深度神经网络仍极易受对抗扰动影响,对抗训练(AT)是提升鲁棒性的常用方法。但平均鲁棒精度的提升常掩盖显著的类别差异:部分类别鲁棒性增强,其余类别在攻击下仍可能极度脆弱。这种失衡引发了重要的对抗公平性问题,尤其在视觉任务中,要求所有类别具备可靠鲁棒性。为解决该挑战,我们提出RL-FAT,一种受强化学习启发的公平对抗训练框架,其利用对抗预测的基于策略梯度的反馈。RL-FAT将预测分布视为策略,结合基于正确性的预测奖励与类别价值估计,计算用于策略梯度优化的类别特定优势,使模型自适应聚焦类别误分类。此外,我们引入公平强调对抗损失,为高对抗损失类别分配更强训练压力,从而缓解类别鲁棒性失衡。通过结合强化驱动的自适应与公平强调正则化,RL-FAT在提升对抗鲁棒性的同时,促进类别间更均衡的鲁棒性分布。大量实验表明,与标准对抗训练基线相比,我们的方法达到了有竞争力的鲁棒精度,并大幅降低了类别鲁棒性失衡。
英文摘要:
Deep neural networks remain highly vulnerable to adversarial perturbations, and adversarial training (AT) has become a widely used approach for improving robustness. However, improvements in average robust accuracy often mask substantial class-wise disparities: while some classes become more robust, others may remain disproportionately vulnerable under attack. This imbalance raises an important adversarial fairness concern, particularly in vision tasks where reliable robustness is expected across all categories. To address this challenge, we propose \textbf{RL-FAT}, a reinforcement-learning-inspired fair adversarial training framework that uses policy-gradient based feedback from adversarial predictions. RL-FAT interprets the prediction distribution as a policy and combines correctness-based prediction rewards with class-wise value estimates to compute class-specific advantages for policy-gradient optimization. This enables the model to adaptively focus on class-wise misclassification. Furthermore, we introduce a fairness-emphasis adversarial loss that assigns stronger training pressure to classes with high adversarial loss, thereby mitigating class-wise robustness disparity. By combining reinforcement-driven adaptation with fairness-emphasis regularization, RL-FAT improves adversarial robustness while promoting a more balanced robustness distribution across classes. Extensive experiments demonstrate that our method achieves competitive robust accuracy and substantially reduces class-wise robustness imbalance compared with standard adversarial training baselines.