DiffSynth-Music:用于可控音乐生成的音频条件KV缓存适配器
DiffSynth-Music: Audio-Conditioned KV-Cache Adapters for Controllable Music Generation
- Alibaba Group(阿里巴巴集团)
- Shanghai Jiao Tong University(上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
针对文本歌词对音乐节奏、旋律控制不足的问题,提出DiffSynth-Music框架,通过音频条件KV缓存注入支持五种控制类型,提升可控性与歌词保真度。
中文摘要 AI 辅助
文本和歌词指定了广泛的音乐特征和演唱内容,但对音乐节奏、旋律和基于参考的风格提供了有限的控制。我们提出了DiffSynth-Music(此https URL),一个通过逐层键值注入将可组合的音频条件添加到音乐合成骨干框架中的框架。三个模板模型,即Control、Prosody和Reference,从骨干扩散变换器初始化,并使用条件流匹配进行训练。它们支持五种控制类型:节拍、人声、伴奏、韵律和参考音频。共享的变分自编码器将条件波形映射到公共潜在空间,使其注意力记忆能够被组合。当模板时间步固定在干净数据端点且其他输入保持不变时,每个控制缓存被计算一次并在整个采样过程中重用。训练对从音乐录音中通过节拍提取、源分离、人声再合成和参考片段选择获得。在中文和英文歌曲上的单控制评估表明,与骨干模型相比,在所有五种控制类型上的一致性得到改善,并且在人声条件下歌词保真度更高。自动音乐质量和指令遵循得分与所评估的基础模型大致相当,存在特定指标的权衡。我们发布了三个模板模型,以支持可控音乐生成的研究和创意应用。
英文摘要
Text and lyrics specify broad musical characteristics and sung content but offer limited control over musical timing, melody, and reference-based style. We introduce DiffSynth-Music (https://modelscope.cn/models/DiffSynth-Studio/DiffSynth-Music), a framework that adds composable audio conditioning to a music synthesis backbone through layer-wise key-value injection. The three template models, Control, Prosody, and Reference, are initialized from the backbone diffusion transformer and trained with conditional flow matching. They support five control types: beats, vocals, accompaniment, prosody, and reference audio. A shared variational autoencoder maps conditioning waveforms into a common latent space, enabling their attention memories to be combined. With the template timestep fixed at the clean-data endpoint and other inputs held constant, each control cache is computed once and reused throughout sampling. Training pairs are derived from music recordings using beat extraction, source separation, vocal resynthesis, and reference-excerpt selection. Single-control evaluations on Mandarin and English songs demonstrate improved adherence across all five control types and better lyric fidelity under vocal conditioning relative to the backbone. Automatic music-quality and instruction-following scores remain broadly comparable to those of the evaluated base models, with metric-specific trade-offs. We release the three template models to support research and creative applications in controllable music generation.