面向印度语言的自发语音检测与域外合成语音泛化的预训练语音编码器评估
Evaluating Pre-trained Speech Encoders for Spontaneous Speech Detection and Out of Domain Synthetic Speech Generalisation in Indic Languages
AI总结:
本研究针对22种印度语言,评估5种预训练语音编码器的自发语音检测与域外合成语音泛化能力,揭示编码器在语言可区分性与自发性检测间的权衡,为深度伪造检测器训练数据选择提供依据。
AI中文摘要:
基于Transformer的模型在区分自发语音与脚本语音、自然语音与合成语音方面展现出较强的准确率,但这些结果是在一组资源丰富的有限语言基准上取得的,尚未扩展到印度语言,也未利用嵌入几何解释编码器行为或深度伪造泛化失败。为解决这些空白,我们在22种印度语言上评估了5种冻结的Transformer编码器:AST、Vaani-FastConformer、Wav2vec2、Whisper和BEATs,并在4种TTS模型上开展了多系统TTS泛化实验。除准确率外,我们还进行了语言隔离探测和质心邻近度分析。探测显示存在编码器依赖的语言可区分性与自发性检测之间的权衡;质心分析表明,域外泛化由训练系统与未见TTS嵌入的邻近度预测,而非与自然语音的距离,该发现对实际深度伪造检测器的训练数据选择具有直接意义。
英文摘要:
Transformer-based models have shown strong accuracy in distinguishing spontaneous from scripted speech and natural from synthetic speech, but these results are established on a narrow set of well-resourced language benchmarks and have not been extended across Indic languages, nor has embedding geometry been used to explain encoder behaviour or deepfake generalisation failure. We address these gaps by evaluating five frozen transformer encoders, AST, Vaani-FastConformer, Wav2vec2, Whisper and BEATs, across 22 Indic languages, and by conducting a multi-system TTS generalisation experiment across four TTS models. Beyond accuracy, we present language isolation probing and centroid proximity analysis. Probing reveals an encoder-dependent trade-off between language-discriminability and spontaneity detection. Centroid analysis shows that out-of-domain generalisation is predicted by a training system's proximity to unseen TTS embeddings, not its distance from natural speech, a finding with direct implications for training data selection in real-world deepfake detectors.