将自动音乐混音重新思考为逐声部序列混合
Rethinking Automatic Music Mixing as Sequential Stem Blending
AI总结:
本研究将自动音乐混音重新定义为逐声部序列混合任务,提出潜在流匹配模型并结合退化数据合成策略,在相关基准上验证了方法有效性。
AI中文摘要:
自动音乐混音是将单个音频轨道自动组合成连贯混音的任务,通常采用并行架构处理所有输入轨道。受人类混音工程师逐次处理声部(stem)的启发,本研究提出范式转变,探究能否将自动音乐混音重新表述为逐声部序列混合任务,即每个声部被混合到不断扩展的子混音中。具体而言,我们训练了一个以子混音上下文为条件的潜在流匹配(latent flow matching)模型,可对任意数量的输入轨道进行序列处理。为训练该模型,我们引入基于退化的数据合成策略,利用现有多轨及声源分离数据集模拟真实的声部混合场景。在声部混合和自动音乐混音基准上的实验结果,证明了所提方法的有效性,相关音频示例可在配套演示页面查看。
英文摘要:
Automatic music mixing, the task of automatically combining individual audio tracks into a cohesive mixture, is typically addressed by parallelized architectures that process all input tracks in a single pass. In this work, inspired by how human mix engineers process stems one at a time, we propose a paradigm shift and ask whether automatic music mixing can be reformulated as a sequential stem blending task, where each stem is blended into a growing submix. Specifically, we train a latent flow matching model conditioned on the submix context, enabling sequential processing of an arbitrary number of input tracks. To train the model, we introduce a degradation-based data synthesis strategy that simulates realistic stem blending scenarios from existing multitrack and source separation datasets. Experimental results on both stem blending and automatic music mixing benchmarks demonstrate the effectiveness of the proposed approach. We provide audio examples on the accompanying demo page\footnote{https://sequential-mixing-demo.vercel.app/}.