发表机构
Alibaba Group(阿里巴巴集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对音乐生成难题,提出万松方法,它是基于扩散的模型,可直接生成高保真多语言长歌曲并输出双声道,通过步长蒸馏加快推理,为微调定制提供途径,助力下游编辑任务。
AI 中文摘要
音乐生成基础模型近来备受行业关注。但要实现高效生成、高保真长音频并支持可控性仍具挑战。为此提出万松,一种用于长格式商业级歌曲生成的简单却强大的方法。它是纯基于扩散的模型,能直接生成长达5分钟的高保真多语言歌曲,单次运行输出双声道。其扩散框架通过步长蒸馏实现更快推理,还为微调与定制提供有效途径以支持下游编辑任务。
英文摘要
Music generation foundation models have recently attracted significant industry attention. However, achieving efficient generation and high-fidelity long-form audio while supporting controllability remains challenging. To address these needs, we present \textbf{WanSong}, a simple yet powerful approach for long-form, commercial-grade song generation. Unlike autoregressive (AR) and cascaded multi-stage pipelines (\eg, AR followed by diffusion), \textbf{WanSong} is a pure diffusion-based model that directly generates high-fidelity, multilingual songs up to 5 minutes and outputs dual stems (vocals and background music) in a single run. In addition, our diffusion framework enables faster inference through step-distillation, and offers an efficient pathway for fine-tuning and customization to support downstream editing tasks.
CommentsWan Team, Alibaba Group