发表机构
Sharif University of Technology; University of Tehran; New Uzbekistan University; Okinawa Institute of Science and Technology(谢里夫理工大学; 德黑兰大学; 新乌兹别克斯坦大学; 冲绳科学技术大学院大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对语音基础模型,提出层级稳定化方法,在不需下游对抗样本的情况下提升可迁移鲁棒性,在12个骨干-任务对上显著提高鲁棒准确率。
AI 中文摘要
冻结的语音基础模型(SFMs)使下游适配高效:骨干网络可以保持固定,而任务仅需学习层融合和轻量级分类器。完全对抗微调是获得鲁棒性的标准途径,但为每个任务生成对抗样本并更新骨干网络会牺牲这种效率。我们探讨是否可以在未来任务未知之前就学习鲁棒性。对于冻结的骨干网络和线性分类器,鲁棒性可以通过表征稳定性与决策边界间隔之间的相互作用来理解。这直接引导了我们的设计:我们在隐藏层上稳定表征,而非仅限最后一层,同时保留干净表征;在干净适配选择层混合后,我们保持其固定,仅扩大分类器间隔,无需下游对抗样本。我们在四个任务上评估了Wav2Vec2、HuBERT和WavLM Large,在自适应30 dB攻击下。在12个骨干-任务对中,层级鲁棒化将鲁棒准确率提高了46.4个百分点,而间隔细化在牺牲1.1个百分点干净准确率的情况下额外提高了4.0个百分点。代码和配置可在该https URL获取。
英文摘要
Frozen speech foundation models (SFMs) make downstream adaptation efficient: the backbone can stay fixed while a task learns layer fusion and a lightweight classifier. Full adversarial fine-tuning is a standard route to robustness, but generating adversarial examples and updating the backbone for every task sacrifices that efficiency. We ask whether robustness can instead be learned before future tasks are known. For a frozen backbone and linear classifier, robustness can be understood through the interaction between representation stability and decision-boundary margin. This leads directly to our design: we stabilize representations across the hidden layers, rather than only the final layer, while preserving clean representations; after clean adaptation selects the layer mixture, we keep it fixed and enlarge only the classifier margin, without downstream adversarial examples. We evaluate Wav2Vec2, HuBERT, and WavLM Large on four tasks under adaptive 30 dB attacks. Across 12 backbone-task pairs, hierarchical robustification improves robust accuracy by 46.4 pp, while margin refinement adds 4.0 pp for 1.1 pp of clean accuracy. Code and configurations are available at https://github.com/arefmousavi/hierarchical-robust-sfm.
Comments5 pages, 2 figures, 2 tables. Submitted to IEEE ICASSP 2027