发表机构
Idiap Research Institute; École Polytechnique Fédérale de Lausanne (EPFL); Sharif University of Technology; University of Isfahan(Idiap研究所; 洛桑联邦理工学院; 谢里夫理工大学; 伊斯法罕大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对微调ASR模型易受对抗迁移攻击的问题,提出TransferBreaker统一微调框架,通过基础对抗微调、潜在雅可比正则化和混合梯度对抗微调抑制迁移,在三种语言和四个大模型上将对抗WER从92.6降至27.8。
AI 中文摘要
许多组织对公开可用的预训练自动语音识别(ASR)模型进行微调,并将其部署在黑盒环境中,假设有限的访问权限能提供保护。我们表明这一假设是脆弱的:在公共基础模型上制作的对抗扰动能够有效地迁移到微调后的目标模型,严重降低性能,并对安全关键应用构成担忧。我们提出了TransferBreaker,一个统一的微调框架,通过整合基础对抗微调(将对抗训练限制在基础有效扰动上)、潜在雅可比正则化(通过抑制对抗敏感方向来强制潜在空间不变性)以及HybridGrad-AFT(通过插值来自基础梯度和目标梯度的可迁移扰动来提高对自适应攻击的鲁棒性)来抑制对抗迁移。我们从理论上证明了所有组件,并在三种语言和四个大型ASR模型上评估了TransferBreaker,将对抗词错误率(WER)从92.6降至27.8。我们的代码在此https URL公开可用。
英文摘要
Many organizations fine-tune publicly available pretrained Automatic Speech Recognition (ASR) models and deploy them in black-box settings, assuming limited access provides protection. We show this assumption is fragile: adversarial perturbations crafted on the public base model transfer effectively to fine-tuned target models, severely degrading performance and posing concerns for safety-critical applications. We propose TransferBreaker, a unified fine-tuning framework that suppresses adversarial transfer by integrating Base Adversarial Fine-Tuning, which restricts adversarial training to base-effective perturbations; Latent Jacobian Regularization, which enforces latent-space invariance by suppressing adversarially sensitive directions; and HybridGrad-AFT, which improves robustness against adaptive attacks by interpolating transferable perturbations from base and target gradients. We theoretically justify all components and evaluate TransferBreaker across three languages and four large ASR models, reducing adversarial WER from 92.6 to 27.8. Our code is publicly available at https://github.com/rohban-lab/TransferBreaker.
Comments39 pages, 5 figures