arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.25546cs.SDcs.LGeess.AS

使用MIDI Span条件化的标量量化潜变量进行多乐器音频混合的合成与编辑

Synthesis and editing of multi-instrument audio mixtures using scalar-quantised latents with MIDI Span conditioning

Sungkyun Chang, Keshav Bhandari, Simon Dixon, Emmanouil Benetos

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出SpanSynth-Edit,一种基于流匹配的MIDI引导多乐器音频合成与编辑模型,利用标量量化潜变量和MIDI Span条件化,实现帧内起始控制,并在基准上取得竞争力表现。

中文摘要 AI 辅助

音乐创作通常涉及迭代式精炼,即在保留其余部分的同时修改选定的音乐细节。为支持此类精炼,我们引入了SpanSynth-Edit,一种基于流匹配的模型,用于使用低帧率标量量化潜变量进行MIDI引导的多乐器音频混合的合成与编辑。MIDI Span将乐器标注的音符生命周期编码为具有连续值属性的无序事件集,并将每个集合池化为每个音频潜变量帧的一个条件向量。该模型使用上下文音频进行特定乐器的音色引导,并支持通过从修订后的MIDI重新合成目标区域来进行编辑。在单乐器和多乐器基准上的实验显示了具有竞争力的性能,并展示了帧内起始控制。我们还讨论了基于转录的音符一致性评估的局限性。

英文摘要

Music creation often involves iterative refinement, changing selected musical details while retaining the rest. To support such refinement, we introduce SpanSynth-Edit, a flow-matching model for MIDI-guided synthesis and editing of multi-instrument audio mixtures using low-frame-rate scalar-quantised latents. MIDI Span encodes instrument-labelled note lifecycles as unordered event sets with continuous-valued attributes and pools each set into one conditioning vector per audio-latent frame. The model uses contextual audio for instrument-specific timbre guidance and supports editing by resynthesising the target region from revised MIDI. Experiments on single- and multi-instrument benchmarks show competitive performance and demonstrate within-frame onset control. We also discuss limitations of transcription-based note-adherence evaluation.

发表机构

  • Queen Mary University of London(伦敦玛丽女王大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑