arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

将基于扩散的音乐合成模型应用于人类语音转换

Adapting Diffusion-Based Music Synthesis to Speech and Singing Voice Conversion

Ben Maman, Frank Zalkow, Hans-Ulrich Berendes, Paolo Sani, Christian Dittmar, Meinard Müller

arXiv 2607.13278首次发表:更新:

AI 中文总结

研究将基于扩散的音乐合成模型用于语音转换,通过扩展条件设定、重新解释音色条件等方法,使模型在自然度等方面表现良好,还发现纳入乐器数据的问题及现成特征提取器的作用,凸显跨域模型转移对统一音频生成系统的潜力。

AI 中文摘要

近期基于扩散的生成模型在特定领域音频生成任务中成果显著,但通常较为专门化,对混合或中间音频类型泛化性不佳。本文将原本用于多乐器音乐合成的基于扩散的模型应用于语音转换,在统一框架内涵盖语音和歌唱。具体而言,扩展基于音符的条件设定,纳入语音后验图和音高轮廓,通过特征-wise线性调制将音色条件重新解释为说话者或歌手身份。实验表明,该适配模型在自然度和表演者相似度方面与专用语音转换系统相当或更优,同时保持跨语音和歌唱的精确音高控制。研究还发现纳入乐器训练数据时存在语音保真度限制和音质下降问题,且现成特征提取器能提供有效条件信号,实现无需人工标注的大规模自监督训练。这些结果凸显了跨域模型转移对统一音频生成系统的潜力。

英文摘要

Recent diffusion-based generative models have achieved strong results in domain-specific audio generation tasks such as speech, singing, and instrumental music synthesis. However, these models are typically specialized and do not generalize well to mixed or intermediate audio types. In this work, we adapt a diffusion-based model originally designed for multi-instrument music synthesis to voice conversion, covering both speech and singing within a unified framework. Specifically, we extend musical note-based conditioning to include phonetic posteriorgrams (PPGs) and pitch contours, and reinterpret timbre conditioning as speaker or singer identity via feature-wise linear modulation. Experiments show that the adapted model matches or surpasses a dedicated voice conversion system in terms of naturalness and performer similarity, while maintaining accurate pitch control across speech and singing. At the same time, we observe limitations in phonetic fidelity and a degradation in vocal quality when incorporating instrumental training data. Furthermore, we demonstrate that off-the-shelf feature extractors provide effective conditioning signals, enabling large-scale self-supervised training without manual annotations. These results highlight the potential of cross-domain model transfer towards unified audio generation systems capable of handling speech, singing, and music. Qualitative samples can be found on our project page: https://benadar293.github.io/voice-conversion

CommentsAccepted to International Conference on Digital Audio Effects (DAFx) 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑