arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.24770eess.AScs.LGcs.SDeess.SP

XSQ-AST:一种用于定位合成语音伪影的可解释音频频谱图Transformer框架

XSQ-AST: An Explainable Audio Spectrogram Transformer Framework for Localising Synthetic Speech Artifacts

发表机构英国约克大学物理工程技术学院音频实验室 · 英国约克大学计算机科学系 · 瑞典艺电SEED部门
另 1 家 · 查看机构详情
  • AudioLab, School of Physics, Engineering and Technology, University of York, United Kingdom(英国约克大学物理工程技术学院音频实验室)
  • Department of Computer Science, University of York, United Kingdom(英国约克大学计算机科学系)
  • SEED -- Electronic Arts (EA), Sweden(瑞典艺电SEED部门)
  • Electronic Arts (EA), United States(美国艺电)

机构由 AI 辅助整理,请以论文原文为准。

Ben Heritage, Luca Resti, Mónica Villanueva Aylagas, Timothy Mehlenbacher, Konrad Tollmar, James Alfred Walker

首次发表
浏览论文内容

中文总结 AI 辅助

XSQ-AST结合SQ-AST、WhisperX和多种显著性方法,无需重训练即可定位合成语音伪影,经40人听力测试和AUC-ROC验证有效。

中文摘要 AI 辅助

定位合成语音中的伪影仍然具有挑战性,因为大多数评估方法仅产生全局质量分数。本文提出了XSQ-AST,一个结合SQ-AST语音质量模型、WhisperX音素对齐和多种显著性方法的框架,无需模型重训练即可生成时间定位的伪影诊断。通过核密度估计将显著性图投影到连续分布上,并通过音素离散化显著性图投影到音素边界上。一项40名参与者的听力测试在五个感知维度上验证了该框架。Attention Rollout、Attention Flow和一种改编的GradCAM产生的时间分布与听众高亮相关,不同方法最适合不同类型的伪影。AUC-ROC分析确认了区分能力优于随机水平。

英文摘要

Localising artifacts in synthetic speech remains challenging, as most evaluation methods yield only global quality scores. This paper presents XSQ-AST, a framework that combines the SQ-AST speech quality model with WhisperX phoneme alignment and multiple saliency methods to produce temporally localised artifact diagnostics without model retraining. Saliency maps are projected onto continuous distributions via kernel density estimation and onto phoneme boundaries via phoneme-discretised saliency maps. A 40-participant listening test validated the framework across five perceptual dimensions. Attention Rollout, Attention Flow and an adapted GradCAM produced temporal distributions that correlated with listener highlights, with different methods best suited to different artifact types. An AUC-ROC analysis confirmed discrimination above chance.

补充信息

↑