发表机构
Carnegie Mellon University(卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出轻量卷积探针引导Stable Audio Open扩散式音乐生成的方法,无需修改模型,实验显示其旋律连贯性较基线提升2.4倍,证实扩散式音乐潜在空间的音乐结构可恢复且可引导。
AI 中文摘要
近期可控音乐生成的研究多聚焦于自回归模型,使得扩散式系统的相关探索相对不足。本文提出一种轻量方法,用于引导音乐合成潜在扩散模型Stable Audio Open生成音频的音高内容:通过配对音频与MIDI数据,训练一个约含12.5万参数的小型卷积探针,以从模型的变分自编码器潜在空间解码帧级音高类激活。推理阶段,该冻结的探针作为可微损失函数,利用其对去噪潜在变量的梯度,将生成过程推向用户指定的音高类序列,无需对基础模型进行重新训练或架构修改。在覆盖9个文本提示和3个目标旋律的27次评估试验中,经探针引导的生成旋律连贯性较无引导基线提升2.4倍(p<1e-5,Wilcoxon符号秩检验),证明扩散式音乐潜在空间中兼具音乐意义的结构既具备可恢复性,也可被引导。
英文摘要
Recent work on controllable music generation has focused on autoregressive models, leaving diffusion-based systems comparatively underexplored. We present a lightweight method for steering the pitch content of audio produced by Stable Audio Open, a latent diffusion model for music synthesis. A small convolutional probe containing approximately 125k parameters is trained to decode frame-level pitch-class activations from the model's variational autoencoder latent space, using paired audio and MIDI data. At inference time, the frozen probe serves as a differentiable loss function: its gradient with respect to the denoising latent is used to nudge generation toward a user-specified pitch-class sequence, requiring no retraining or architectural modification of the base model. Across 27 evaluation trials spanning 9 text prompts and 3 target melodies, probe-guided generation increases melodic coherence by 2.4x over the unguided baseline (p < 1e-5, Wilcoxon signed-rank test), demonstrating that musically meaningful structure is both recoverable and steerable in diffusion-based music latent spaces.
CommentsAccepted at IEEE MLSP 2026. 6 pages, 3 figures, 3 tables