合成语音归因的原型网络方法
Synthetic Speech Attribution via Prototypical Networks
- Politecnico di Milano(米兰理工大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
本研究提出基于原型网络的合成语音归因方法,在闭集、跨语言和开集条件下达到竞争性性能,同时提供可解释的基于示例的决策依据。
中文摘要 AI 辅助
合成语音归因旨在识别生成语音信号的生成系统,但当前方法通常依赖黑盒神经网络,对其决策提供的洞察有限。本研究探索基于原型的网络作为一种可解释的替代方案,其预测基于与代表性训练样本的比较。我们将ProtoPNet适配到基于频谱图的语音表示,并在MLAAD数据集上于闭集、跨语言和开集条件下评估所提出的框架。实验表明,与基线相比,基于原型的推理在归因性能上达到竞争性或改进水平,同时支持基于示例的解释。这些结果凸显了在合成语音归因中,通过基于原型的建模可以同时实现可解释性和性能。
英文摘要
Synthetic speech attribution aims to identify the generative system responsible for a speech signal, but current approaches typically rely on black-box neural networks that provide limited insight into their decisions. This work investigates prototype-based networks as an interpretable alternative, where predictions are grounded in comparisons with representative training examples. We adapt ProtoPNet to spectrogram-based speech representations and evaluate the proposed framework on the MLAAD dataset under closed-set, cross-lingual, and open-set conditions. Experiments show that prototype-based reasoning achieves competitive or improved attribution performance compared with the baseline while enabling example-based explanations. These results highlight that interpretability and performance can be jointly achieved in synthetic speech attribution through prototype-based modeling.