CookVoice:风格可控的多模态人声生成统一框架
CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation
浏览论文内容
中文总结 AI 辅助
CookVoice是统一的多模态人声生成框架,分解人声为内容、韵律、风格,支持多任务,参数少、推理高效,可控性强且生成质量接近基线模型。
中文摘要 AI 辅助
人声生成在语音生成、歌唱语音生成、语音克隆和语音编辑领域已取得快速进展。然而,现有大多数系统针对特定任务设计,常依赖任务相关的架构、控制信号或自回归解码,限制了细粒度可控性与推理效率。本文提出CookVoice,一种用于多模态、多风格、多任务人声生成的统一框架。CookVoice将人声分解为内容、韵律、风格三个关键因素,可在统一模型内实现语音与歌唱语音生成。为实现精确灵活的可控性,设计了一种灵活对齐策略,将文本、风格与韵律控制信号映射到频谱图的帧级别,该设计使CookVoice支持文本转语音、文本转歌唱语音、风格可控生成、语音模仿、语音转换及语音编辑等多种任务。实验结果表明,CookVoice的生成质量可与现有文本转语音、文本转歌唱语音的基线模型相当,同时具备更强的风格与韵律可控性;此外,CookVoice仅用4351万参数,且推理仅需最少4个ODE步骤,即可达到与大规模基线模型相当的性能,是适用于现实世界人声生成应用的实用解决方案。演示页面可访问此https URL。
英文摘要
Human voice generation has made rapid progress in speech generation, singing voice generation, voice cloning, and voice editing. However, most existing systems are designed for specific tasks and often rely on task-dependent architectures, control signals, or autoregressive decoding, limiting fine-grained controllability and inference efficiency. In this paper, we propose CookVoice, a unified framework for multimodal, multi-style, and multi-task human voice generation. CookVoice decomposes the human voice into three key factors: content, prosody, and style, enabling both speech and singing voice generation within a unified model. To achieve precise and flexible controllability, we design a flexible alignment strategy that maps text, style, and prosody control signals onto the frame-level of spectrogram. This design allows CookVoice to support a wide range of tasks, including text-to-speech, text-to-singing voice, style-controllable generation, voice mimicry, voice conversion, and voice editing. Experimental results show that CookVoice achieves generation quality comparable to existing Text-to-Speech and text-to-singing voice baselines, while providing stronger style and prosody controllability. Moreover, CookVoice achieves comparable performance to large-scale baselines with only 43.51 million parameters and efficient inference using as few as 4 ODE steps, making it a practical solution for real-world human voice generation applications. Demo page is available at https://haoweilou.github.io/CookVoice/.
发表机构
- UNSW Sydney(新南威尔士大学悉尼分校)
- Dolby Laboratories(杜比实验室)
机构由 AI 辅助整理,请以论文原文为准。