AI 中文总结
研究虚假音频检测问题,提出逐层决策融合方法,利用大型语音模型多层表示,在每层决策后融合,相比其他基线在自然数据集上性能最佳,且模型更透明利于分析决策机制。
AI 中文摘要
近期的虚假音频检测方法常利用大型语音模型来获得强大的语音表示。这些模型通常非常深,能提供多层表示。然而,当前工作常仅依赖单层表示或特征融合来提取一个话语级表示进行决策。这些方法可能未充分利用多层的丰富信息且可能导致特征崩溃。我们提出一种新颖的逐层决策融合方法,在每层决策后进行融合,与其他强基线相比,在自然数据集上实现了最佳跨数据集性能(EER 6.90%)。我们的模型设计还使模型更透明,能进行详细分析以揭示决策的潜在机制。
英文摘要
Recent fake audio detection methods often leverage large speech models to achieve robust speech representations. These models are typically very deep, providing multiple layer-wise representations. However, current works often rely solely on single layer representation or feature fusion to extract one utterance-level representation for decision making. These methods risk underutilizing rich information from multiple layers and might induce feature collapse. We propose a novel layer-wise decision fusion method that applies fusion after per-layer decision making and achieves the best cross-dataset performance on In-the-Wild dataset (EER 6.90%) compared to other strong baselines. Our model design also makes the model more transparent, allowing us to conduct detailed analysis to reveal the underlying mechanism of decision making.
CommentsAccepted to Interspeech 2025
DOI:10.21437/Interspeech.2025-1543