发表机构
Jinan University(暨南大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出REIMU,通过对比不同架构在ASVspoof数据集上的表现,发现异构算子分配的设计可在减少下游参数的同时保持语音深度伪造检测的竞争力,为参数高效检测提供了新方向。
AI 中文摘要
文本转语音和语音转换系统生成的语音逼真度不断提高,对媒体完整性和语音认证构成了日益严峻的挑战。自监督学习(SSL)已大幅推进语音深度伪造检测,其中下游骨干网络通常通过单次前向传播处理SSL表示。本研究针对该任务探究循环分层推理的实际效果,将该受控研究命名为REIMU,在四个基础规模SSL前端上系统对比传统单次前向骨干网络、权重共享循环、同构HRM及异构HRM,进一步研究结合自注意力与线性注意力的异构高低层级模块。在ASVspoof 2019和2021评估集上的实验表明,循环与分层分解本身无法提升检测性能,而异构算子分配能提供更具竞争力的配置;值得注意的是,异构设计相较匹配基线使用10.8%更少的下游参数仍保持竞争力,展现出其在参数高效型语音深度伪造检测中的潜力。
英文摘要
The increasing realism of speech generated by text-to-speech and voice conversion systems poses growing challenges to media integrity and voice authentication. Self-supervised learning (SSL) has substantially advanced speech deepfake detection, where downstream backbones conventionally process SSL representations through a single forward pass. This work investigates the practical effectiveness of recurrent hierarchical reasoning for this task. We term this controlled study REIMU and systematically compare conventional single-pass backbones, weight-shared recurrence, homogeneous HRM, and heterogeneous HRM across four Base-scale SSL frontends. We further examine heterogeneous high- and low-level modules that combine self-attention with linear attention. Experiments on the ASVspoof 2019 and 2021 evaluation sets show that recurrence and hierarchical decomposition do not inherently improve detection, whereas heterogeneous operator assignment provides a more competitive configuration. Notably, the heterogeneous design remains competitive while using 10.8\% fewer downstream parameters than the matched baseline, demonstrating its potential for parameter-efficient speech deepfake detection.