AI 中文总结
本文通过三项受控听觉实验分析自动混音,探究两阶段系统性能提升的来源,发现不当分组会降低下游性能,两阶段变体优于单阶段基线,支持局部与全局混音分离的设计原则。
AI 中文摘要
自动混音将多轨录音转换为感知上连贯、平衡且审美一致的混音。在实际制作中,由于音轨数量多、乐器种类多样以及音轨间存在强依赖关系,该任务颇具挑战性。两阶段系统通过将组内处理与组间混音分离来应对这种复杂性,但目前尚不清楚其性能提升源于更强的组件模型还是明确的任务分解。我们通过三项受控听觉实验对自动混音开展面向子任务的分析,研究全混音模型是否可迁移至组内混音、下游模型是否可补偿分组与响度误差,以及两阶段分解是否能提升全混音质量。在三首密集的流行乐与摇滚乐片段中,评估模型的迁移效果存在差异;不当分组会导致下游性能明显下降,而改变响度关系的影响较弱且取决于模型。两种两阶段变体均显著优于其对应的单阶段基线。这些发现支持将局部平衡与全局混音协调的明确分离作为自动混音的实用设计原则。代码与音频示例可在线获取。
英文摘要
Automatic mixing transforms multitrack recordings into perceptually coherent, balanced, and aesthetically consistent mixes. In real-world production, this task is challenging due to large track counts, diverse instrumentation, and strong inter-track dependencies. Two-stage systems address this complexity by separating intra-group processing from inter-group mixing, yet it remains unclear whether their gains arise from stronger component models or from explicit task decomposition. We present a subtask-oriented analysis of automatic mixing through three controlled listening experiments. We investigate whether full-mix models transfer to intra-group mixing, whether downstream models compensate for grouping and loudness errors, and whether two-stage decomposition improves full-mix quality. Across three dense pop and rock excerpts, transfer differs between the evaluated models; inappropriate grouping causes clear downstream degradation, while altered loudness relationships have weaker and model-dependent effects. Both two-stage variants significantly outperform their corresponding single-stage baselines. These findings support explicit separation of local balance and global mix coordination as a useful design principle for automatic mixing. Code and audio examples are available online.
CommentsAccepted at the International Society for Music Information Retrieval Conference (ISMIR 2026). 6 pages, 5 figures. https://sparrowreivun.github.io/TwoStageMixingAnalysis/