使用MIDI Span条件化的标量量化潜变量进行多乐器音频混合的合成与编辑
Synthesis and editing of multi-instrument audio mixtures using scalar-quantised latents with MIDI Span conditioning
浏览论文内容
中文总结 AI 辅助
本文提出SpanSynth-Edit,一种基于流匹配的MIDI引导多乐器音频合成与编辑模型,利用标量量化潜变量和MIDI Span条件化,实现帧内起始控制,并在基准上取得竞争力表现。
中文摘要 AI 辅助
音乐创作通常涉及迭代式精炼,即在保留其余部分的同时修改选定的音乐细节。为支持此类精炼,我们引入了SpanSynth-Edit,一种基于流匹配的模型,用于使用低帧率标量量化潜变量进行MIDI引导的多乐器音频混合的合成与编辑。MIDI Span将乐器标注的音符生命周期编码为具有连续值属性的无序事件集,并将每个集合池化为每个音频潜变量帧的一个条件向量。该模型使用上下文音频进行特定乐器的音色引导,并支持通过从修订后的MIDI重新合成目标区域来进行编辑。在单乐器和多乐器基准上的实验显示了具有竞争力的性能,并展示了帧内起始控制。我们还讨论了基于转录的音符一致性评估的局限性。
英文摘要
Music creation often involves iterative refinement, changing selected musical details while retaining the rest. To support such refinement, we introduce SpanSynth-Edit, a flow-matching model for MIDI-guided synthesis and editing of multi-instrument audio mixtures using low-frame-rate scalar-quantised latents. MIDI Span encodes instrument-labelled note lifecycles as unordered event sets with continuous-valued attributes and pools each set into one conditioning vector per audio-latent frame. The model uses contextual audio for instrument-specific timbre guidance and supports editing by resynthesising the target region from revised MIDI. Experiments on single- and multi-instrument benchmarks show competitive performance and demonstrate within-frame onset control. We also discuss limitations of transcription-based note-adherence evaluation.
发表机构
- Queen Mary University of London(伦敦玛丽女王大学)
机构由 AI 辅助整理,请以论文原文为准。