arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

合成语音,真实信号:通过语音克隆实现副语言保留和跨语言增强

Synthetic Speech, Real Signal: Paralinguistic Preservation and Cross-Lingual Augmentation via Voice Cloning

Roseline Polle, Owen Parsons, George Fairs, Luis Miguel San Martin Fernandez, Cole Looney, Xiaoliang Wu, Alexandra Livia Georgescu, Stefano Goria

arXiv 2607.22304首次发表:更新:

发表机构

thymia; The University of Edinburgh; University of Southampton(胸腺公司; 爱丁堡大学; 南安普顿大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对语音合成数据增强在副语言任务应用少的问题,对八个语音克隆模型在五个副语言任务上进行基准测试,还将英语临床语音克隆成日语,发现其在低资源语言临床语音数据增强方面有前景。

AI 中文摘要

语音中的合成数据增强在诸如自动语音识别等语言任务中是常见做法,但在副语言任务中应用较少,尤其是在临床任务中,标记数据成本高且部分患者群体代表性不足。语音克隆是一种增强方法,通常根据语音可懂度(词错误率)或说话者相似度进行评估,而非下游性能,且尚不清楚其是否保留了此类任务所依赖的副语言信号。我们在公共和临床数据集的五个副语言任务上对八个语音克隆模型进行基准测试,发现大多数模型能在信号略有退化的情况下保留信号。然后我们将英语临床语音克隆成日语,发现在克隆数据上训练在检测真实日语语音中的抑郁和焦虑方面优于原始跨语言转移,这表明语音克隆是增强低资源语言临床语音数据的一个有前景的方向。

英文摘要

Synthetic data augmentation in speech is common practice for linguistic tasks like ASR, but has seen far less work for paralinguistic ones, especially clinical tasks where labelled data is expensive and some patient groups are underrepresented. Voice cloning is one such augmentation approach, but is typically evaluated on speech intelligibility (WER) or speaker similarity (SS) rather than on downstream performance, and it remains unclear whether these preserve the paralinguistic signal such tasks depend on. We benchmark eight voice cloning models on five paralinguistic tasks across public and clinical datasets, showing most preserve signal with modest degradation. We then clone English clinical speech into Japanese and find that training on cloned data outperforms raw cross-lingual transfer for depression and anxiety detection on real Japanese speech, suggesting voice cloning is a promising direction for augmenting clinical speech data in low-resource languages.

Comments5 pages, 3 figures. Accepted at Interspeech 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑