发表机构
ByteDance; Zhejiang University(字节跳动; 浙江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SwanTale是支持指令式与零样本任务的多说话人语音音频生成模型,通过SwanData-Caption数据构建与SwanVAE、Unified MoE等技术实现,在相关任务指标中表现最优。
AI 中文摘要
动画配音、广播剧、电影、广告、游戏、播客及短视频制作等场景常需语音与音频生成,创作者可能需要在无参考录音的情况下设计音色、用自然语言控制说话人风格、支持带环境与音效的声学场景,并后续复用设计的音色,因此支持指令式与零样本任务的多说话人语音与音频生成至关重要。指令式任务需要环境、说话人风格及细粒度内容的描述文本,零样本任务则结合参考音频与相同细粒度内容。我们从数据与模型两方面解决这些任务:首先提出SwanData-Caption,其清洗原始语音与音频数据、补充针对性合成覆盖内容,并标注多样且准确的多级描述文本;接着提出SwanTale,这是支持零样本与指令式任务的多说话人高表现力语音与音频生成模型,引入SwanVAE实现高质量多音频模态生成,采用奖励条件质量控制与Engram条件,结合统一混合专家模型(Unified MoE)进行多任务与多音频模态建模,还使用课程学习与GRPO后训练让模型逐步学习并强化能力。实验结果显示,SwanTale在多项关键零样本与指令式指标中领先,在两类任务中均取得最优表现力得分,支持涉及多说话人语音与音频的复杂指令式生成,演示可在指定链接查看#swantale。
英文摘要
Speech and audio generation is often needed in animation dubbing, audio drama, movies, advertising, games, podcasts, and short-video production. In these scenarios, creators may need to design voices without reference recordings, control speaker styles with natural language, support acoustic scenes with environments and audio effects, and later reuse the designed voices. Therefore, it is important to support multi-speaker speech and audio generation for both instruct and zero-shot tasks. The instruct task requires a caption of the environment, speaker styles, and fine-grained content, while the zero-shot task uses reference audio together with the same fine-grained content. We address these tasks from both the data and model sides. First, we propose SwanData-Caption, which cleans raw speech and audio data, adds targeted synthetic coverage, and annotates diverse and accurate multi-level captions. Then, we propose SwanTale, a multi-speaker expressive speech and audio generation model that supports both zero-shot and instruct tasks. We introduce SwanVAE to support high-quality multi-audio-modality generation. Then, we adopt reward-conditioned quality control and Engram conditioning, along with Unified MoE for multi-task and multi-audio-modality modeling. In addition, we use curriculum learning and GRPO post-training to let the model progressively learn and strengthen its capabilities. Experimental results show that SwanTale leads on multiple key zero-shot and instruct metrics, achieves the best expressiveness scores in both tasks, and supports complex instruct generation involving multi-speaker speech and audio. Demos can be found at https://swanaigc.github.io/#swantale.
CommentsTechnical Report by ByteDance