arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VoiceDesigner:基于统一扩散建模与数据增强的文本到语音生成与编辑

VoiceDesigner: Text-to-Voice Generation and Editing via Unified Diffusion Modeling and Data Augmentation

Jiarui Hai, Karan Thakkar, Ke Chen, Yunyun Wang, Jiaqi Su, Rithesh Kumar, Mounya Elhilali, Zeyu Jin

arXiv 2608.13613首次发表:更新:

发表机构

Johns Hopkins University; Adobe Research(约翰斯·霍普金斯大学; 奥多比研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

VoiceDesigner是基于统一扩散建模与数据增强的文本到语音生成编辑框架,通过构建多样化语音数据集和改进的扩散Transformer,实现语音生成与编辑,对齐性与感知质量优于现有模型。

AI 中文摘要

生成模型的近期突破已使文本到语音生成(TTV)成为可能,可直接从文本语音描述合成语音。但现有系统面临两大关键挑战:一是难以生成涵盖真实人类说话者与虚构角色的多样化语音;二是缺乏稳健且灵活的语音编辑能力,如语音克隆及修改情感、语调等属性的能力。本文提出VoiceDesigner,一个支持多样化可控语音设计的文本到语音生成与编辑统一框架。为解决上述挑战,从两方面提出方案:其一,开发混合数据流水线,利用数字信号处理技术与语音生成模型构建涵盖真实与虚构语音的多样化语音数据集;其二,引入经架构改进的扩散Transformer,以更好处理复杂条件并提升多任务性能,实现统一的语音生成与编辑。通过主观与客观评估,VoiceDesigner在语音描述与编辑指令的提示对齐上表现更优,同时与现有最先进TTV模型相比,保持了有竞争力的感知质量与语音可用性。

英文摘要

Recent breakthroughs in generative models have made text-to-voice generation (TTV) possible, enabling the synthesis of speech directly from textual voice descriptions. However, existing systems face two key challenges. First, they struggle to generate a diverse range of voices, spanning real-world human speakers and fictional characters. Second, they lack robust and flexible voice editing capabilities, such as voice cloning and the ability to modify attributes like emotion and tone. In this paper, we propose VoiceDesigner, a unified framework for text-to-voice generation and editing that supports diverse and controllable voice design. To tackle the above challenges, we propose solutions from two perspectives. First, we develop a hybrid data pipeline that leverages digital signal processing techniques and speech generation models to construct a diverse voice dataset covering both real-world and fictional voices. Second, we introduce a diffusion transformer with architectural improvements to better handle complex conditioning and enhance multi-task performance, enabling unified voice generation and editing. Through subjective and objective evaluations, VoiceDesigner achieves superior prompt alignment with both voice descriptions and editing instructions, while maintaining competitive perceptual quality and voice usability compared to state-of-the-art TTV models.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑