发表机构
University of Zurich; Idiap Research Institute; Brno University of Technology(苏黎世大学; Idiap研究所; 布尔诺理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对阿拉伯语音深度伪造检测,提出冻结 W2V-BERT-2.0 骨干并采用小波提示微调(更新<1%参数),结合跨层注意力与注意力统计池化,按轨道分别采用多窗口和信道匹配中心裁剪集成,在 ArA-DF 2026 上取得 1.96% 和 1.04% 的 EER。
AI 中文摘要
对于具有方言多样性的低资源语言,检测合成语音和语音转换语音仍然困难,系统必须泛化到不同区域方言和未见过的声学信道。我们介绍了 UZH-CL 提交至 ArA-DF 2026 共享任务(阿拉伯语音深度伪造检测)的方案,涵盖轨道 1(方言泛化)和轨道 2(声学鲁棒性)。我们冻结 W2V-BERT-2.0 骨干网络,并通过小波提示微调(Wavelet Prompt Tuning)进行适配,仅更新不到 1% 的参数;使用跨层注意力聚合多层表示,并采用通用注意力统计池化头,而非专门的图后端。通过改变适配策略、训练数据、增强方式和编码器家族,我们获得了互补的检测器。我们发现两种偏移类型需要不同的融合机制:方言泛化采用广泛的多窗口集成,声学鲁棒性采用紧凑的、信道匹配的中心裁剪集成。官方评估在轨道 1 上取得 1.96% 的等错误率(第 6 名),在轨道 2 上取得 1.04% 的等错误率(第 3 名),相对于 XLS-R+AASIST 基线分别实现了 87% 和 96% 的相对降低。
英文摘要
Detecting synthetic and voice-converted speech remains difficult for low-resource languages with dialectal diversity, where systems must generalize across regional dialects and unseen acoustic channels. We present the UZH-CL submission to the ArA-DF 2026 Shared Task on Arabic speech deepfake detection, covering Track~1 (dialect generalization) and Track~2 (acoustic robustness). We freeze a W2V-BERT-2.0 backbone and adapt it with \emph{Wavelet Prompt Tuning}, updating under 1% of parameters, and aggregate multi-layer representations with cross-layer attention and a general attentive-statistics pooling head rather than a specialized graph backend. Complementary detectors are obtained by varying adaptation strategy, training data, augmentation, and encoder family. We find that the two shift types require different fusion regimes: a broad multi-window ensemble for dialect generalization, and a compact, channel-matched, center-crop ensemble for acoustic robustness. Official evaluation yields 1.96% EER on Track~1 (6th place) and 1.04% EER on Track~2 (3rd place), corresponding to 87% and 96% relative reductions over the XLS-R+AASIST baselines.
Commentsthis is the submitted system of the UZH-CL team for ArA-DF 2026 challenge