arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39722cs.SD

合成语音归因的原型网络方法

Synthetic Speech Attribution via Prototypical Networks

  • Politecnico di Milano(米兰理工大学)

机构由 AI 辅助整理,请以论文原文为准。

Viola Negroni, Paolo Bestagini, Stefano Tubaro

中文总结 AI 辅助

本研究提出基于原型网络的合成语音归因方法,在闭集、跨语言和开集条件下达到竞争性性能,同时提供可解释的基于示例的决策依据。

中文摘要 AI 辅助

合成语音归因旨在识别生成语音信号的生成系统,但当前方法通常依赖黑盒神经网络,对其决策提供的洞察有限。本研究探索基于原型的网络作为一种可解释的替代方案,其预测基于与代表性训练样本的比较。我们将ProtoPNet适配到基于频谱图的语音表示,并在MLAAD数据集上于闭集、跨语言和开集条件下评估所提出的框架。实验表明,与基线相比,基于原型的推理在归因性能上达到竞争性或改进水平,同时支持基于示例的解释。这些结果凸显了在合成语音归因中,通过基于原型的建模可以同时实现可解释性和性能。

英文摘要

Synthetic speech attribution aims to identify the generative system responsible for a speech signal, but current approaches typically rely on black-box neural networks that provide limited insight into their decisions. This work investigates prototype-based networks as an interpretable alternative, where predictions are grounded in comparisons with representative training examples. We adapt ProtoPNet to spectrogram-based speech representations and evaluate the proposed framework on the MLAAD dataset under closed-set, cross-lingual, and open-set conditions. Experiments show that prototype-based reasoning achieves competitive or improved attribution performance compared with the baseline while enabling example-based explanations. These results highlight that interpretability and performance can be jointly achieved in synthetic speech attribution through prototype-based modeling.

补充信息

↑