发表机构
SHOU; Tongji University(上海海洋大学; 同济大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文重新审视视觉Transformer中基于差分隐私的认证防御,发现补丁嵌入后注入噪声时拉普拉斯机制因灵敏度约束几何形状而崩溃,改用谱范数约束并推导无维度隐私保证以解决失效问题。
AI 中文摘要
融合差分隐私的认证防御已被证明在卷积神经网络(CNNs)上有效,为对抗范数有界攻击者提供了严格的鲁棒性保证。然而,像素差分隐私(PixelDP)的认证鲁棒性行为在目前主导深度学习领域的自注意力架构中仍未得到充分探索。鉴于Transformer从数字世界到物理世界对我们的日常应用产生了深远影响,研究通过差分隐私式稳定性实现的认证鲁棒性至关重要。为填补这一研究空白,我们重新审视了视觉Transformer中的这一构造,并识别出一种在卷积设置中基本隐藏的失效模式。当在补丁嵌入后注入噪声时,具有继承的分组ℓ1灵敏度界的拉普拉斯机制在所有噪声尺度下崩溃至机会水平精度,而高斯机制仍可训练。这一对比隔离了失效的根源:并非注入的噪声本身,而是灵敏度约束的几何形状。我们表明,由继承的Δ1,1投影引起的衰减随层宽度和核大小按随机矩阵尺度C/(√M+√N)增加。将ℓ1型约束替换为谱范数约束可消除跨数据集和架构的崩溃,但产生了一个根本性障碍:修复后的模型不再满足标准拉普拉斯证书所需的灵敏度条件。我们通过推导在ℓ2灵敏度下拉普拉斯机制在隐私损失集中下的无维度(ε, δ)-隐私保证来解决这一不匹配问题。
英文摘要
Certified defenses that incorporate differential privacy have proven effective on Convolutional Neural Networks (CNNs), furnishing rigorous robustness guarantees against norm-bounded adversaries. However, the certified robustness behavior of Pixel Differential Privacy (PixelDP) remains largely unexplored with the self-attention architecture now dominating the deep-learning landscape. Given that the Transformer has a profound impact on our daily applications from the digital world to the physical world, it is crucial to study certified robustness through differential-privacy-style stability. To fill this research gap, we revisit this construction in Vision Transformers and identify a failure mode that is largely hidden in the convolutional setting. When noise is injected after the patch embedding, the Laplace mechanism with the inherited grouped $\ell_1$ sensitivity bound collapses to chance-level accuracy across noise scales, whereas the Gaussian mechanism remains trainable. This contrast isolates the source of failure: not the injected noise itself, but the geometry of the sensitivity constraint. We show that the attenuation induced by the inherited $Δ_{1,1}$ projection increases with layer width and kernel size according to a random-matrix scale $C/(\sqrt{M}+\sqrt{N})$. Replacing the $\ell_1$-type constraint with a spectral-norm constraint eliminates the collapse across datasets and architectures, but creates a fundamental obstacle: the repaired models no longer satisfy the sensitivity condition required by the standard Laplace certificate. We resolve this mismatch by deriving a dimension-free $(\varepsilon,\ δ)$-privacy guarantee for the Laplace mechanism under $\ell_2$ sensitivity through concentration of the privacy loss.