发表机构
School of Computing, Australian National University(澳大利亚国立大学计算机学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出基于极分解的谱梯度正交化方法,可在零额外隐私成本下提升差分隐私训练效果,在大批次高容量视觉模型训练中实现显著准确率提升且降低方差,结合时间去噪可取得最优结果。
AI 中文摘要
差分隐私训练会向裁剪后的梯度添加各向同性高斯噪声,导致每个奇异方向的噪声程度相同。在视觉模型中,空间相关性会将梯度能量集中到低秩子空间,大部分噪声会落在几乎无信号的方向上。本文引入了基于极分解的谱梯度正交化作为后处理步骤,可在零额外隐私成本的情况下,从含噪梯度的低秩结构中恢复方向信号。该方法的效用存在相变:仅当每个方向的谱信噪比(SNR)足以支撑奇异向量恢复时,正交化才能提升准确率;在低SNR场景下,梯度的方向偏差会被近乎随机的正交更新取代,该变换会产生负面影响。恢复阈值由梯度的谱间隙决定,且在大批次规模下可达到该阈值。实验表明,收益随模型容量提升而扩大:在WRN-28-10(批次大小B=4096)上,谱正交化较DP-SGD实现了20.9%的准确率提升,在ResNet-18上则实现了14.9%的提升,同时将运行间方差降低了2至3倍。在微调场景中,谱正交化可达到DP-Adam的稳定性,同时保持一阶内存占用。将谱正交化与时间去噪结合,在epsilon=4的CIFAR-10数据集上取得了50.3%的准确率,为所有测试配置中的最高值。这些增益仅适用于中高SNR场景,如大批次训练高容量模型;小批次或低SNR场景更适合使用DP-SGD或时间去噪方法。
英文摘要
Differentially private training adds isotropic Gaussian noise to clipped gradients, corrupting every singular direction equally. In vision models, where spatial correlation concentrates gradient energy into a low-rank subspace, most of this noise falls in directions that carry little signal. Spectral gradient orthogonalization via polar decomposition is introduced as a post-processing step that recovers directional signal from the noisy gradient's low-rank structure at zero additional privacy cost. A phase transition governs the utility of this approach: orthogonalization improves accuracy only when the per-direction spectral signal-to-noise ratio (SNR) suffices for singular vector recovery; in low-SNR regimes, the directional bias of the gradient is replaced by a nearly random orthogonal update, and the transformation is harmful. The recovery threshold is determined by the spectral gap of the gradient and is surpassed at large batch sizes. Empirically, the benefit scales with model capacity: spectral orthogonalization achieves a +20.9% improvement over DP-SGD on WRN-28-10 (B = 4096) and +14.9% on ResNet-18, while reducing inter-run variance by a factor of two to three. In the fine-tuning regime, spectral orthogonalization matches the stability of DP-Adam while maintaining a first-order memory footprint. Combining spectral with temporal denoising yields 50.3% on CIFAR-10 (epsilon = 4), the highest accuracy in any tested configuration. These gains are specific to moderate-to-high-SNR regimes such as large-batch training of higher-capacity models. Small-batch or low-SNR settings are better served by DP-SGD or temporal denoising.
CommentsAccepted at ECCV 2026