训练动态中鲁棒性早期出现但无法保留
Robustness Emerges Early in Training Dynamics, but Is Not Preserved
- Institute of Microelectronics, Chinese Academy of Sciences(中国科学院微电子研究所)
- University of Chinese Academy of Sciences(中国科学院大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对深度神经网络训练中早期出现的鲁棒性无法保留的问题,提出含EPS和AWR两种无参数策略的框架,在多任务中显著提升性能。
AI中文摘要:
对自然干扰的鲁棒性仍是深度神经网络的核心挑战。本文中,我们发现一种鲁棒性衰退现象:浅层在训练早期会自发形成鲁棒表示和平滑损失景观,但这些特性在标准收敛过程中无法保留。为解决该问题,我们提出一个框架,通过对训练动态进行策略性干预来稳定经验识别的早期出现的鲁棒先验。该方法包含两种无参数策略:早期阶段稳定(EPS)和非对称权重还原(AWR),可在不修改模型架构或引入可学习参数的情况下稳定或恢复鲁棒的浅层配置。大量实验在各种基准和架构上验证了该框架的有效性,在下游迁移、动态适应及多种计算机视觉应用中取得显著提升。
英文摘要:
Robustness to natural corruptions remains a fundamental challenge for deep neural networks. In this paper, we identify a robustness fading phenomenon where shallow layers spontaneously develop robust representations and flat loss landscapes in early training, yet these properties are not preserved during standard convergence. To address this, we propose a framework that performs strategic interventions on training dynamics to stabilize the empirically identified early-emergent robust priors. Our approach includes two parameter-free strategies: Early-Phase Stabilization~(EPS) and Asymmetric Weight Reversion~(AWR), which stabilize or recover robust shallow configurations without modifying the model architecture or introducing learnable parameters. Extensive experiments demonstrate the efficacy of our framework across various benchmarks and architectures, yielding significant gains in downstream transfer, dynamic adaptation, and diverse computer vision applications.