发表机构
Indian Institute of Management Bangalore; Indian Statistical Institute, Kolkata(印度管理学院班加罗尔分校; 印度统计研究所加尔各答分所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对神经分类器参数不可识别导致鲁棒训练统计基础不足的问题,提出无需可识别性假设的S散度训练,经验极小值收敛至总体最优等价类,在基准数据集上性能与现有鲁棒方法相当。
AI 中文摘要
通过交叉熵最小化训练的神经网络分类器对标签噪声和对抗污染高度敏感。尽管鲁棒替代方案能提供有界影响并抵抗损坏,但它们在深度学习场景下的统计基础尚不充分,原因在于一个根本难点:神经参数化是不可识别的,因此总体损失极小值是参数的等价类,而非唯一的点。我们基于S散度族开发了鲁棒神经分类器的一致性理论,该理论无需可识别性假设。将训练视为在不可识别参数空间上的随机优化,我们证明在温和正则条件下,经验S散度极小值收敛到总体最优等价类,并验证了三种架构选择下的这些条件。我们进一步确定鲁棒训练算法的极限点是经验目标的平稳点。在视觉和语言基准数据集上的实验证实,S散度训练在保持干净数据准确率的同时,表现出与现有鲁棒方法相当的性能。
英文摘要
Neural network classifiers trained by cross-entropy minimization are highly sensitive to label noise and adversarial contamination. While robust alternatives offer bounded influence and resistance to corruption, their statistical foundations in the deep learning setting are insufficient due to a fundamental difficulty: neural parameterizations are non-identifiable, so the population loss minimizer is an equivalence class of parameters, not a unique point. We develop a consistency theory for robust neural classifiers based on the S-divergence family that requires no identifiability assumption. Casting training as stochastic optimization over a non-identifiable parameter space, we prove that empirical S-divergence minimizers converge to the population-optimal equivalence class under mild regularity conditions, and verify these conditions for three architecture choices. We further establish that limit points of the robust training algorithm are stationary points of the empirical objective. Experiments on vision and language benchmark datasets confirm that S-divergence training maintains clean-data accuracy while exhibiting performance competitive with existing robust methods.