发表机构
Queen Mary University of London(伦敦玛丽女王大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出Diff2Mix系统,结合扩散模型与可微分音频效果,实现兼具高质量与可控风格的自动音乐混音,通过多维度评估验证其性能。
AI 中文摘要
自动音乐混音旨在将多轨录音组合成平衡且连贯的音乐作品。由于不同歌曲的内容与混音工程师的主观偏好共同决定最终结果,实用系统需在实现平衡混音的同时支持可控风格变化。然而,多数现有方法将自动混音与混音风格控制视为独立任务,导致单一系统难以同时产出高质量混音并具备可编辑性与风格感知能力。为解决此局限,本文提出Diff2Mix,一种基于扩散模型与可微分混音控制台的生成式自动混音系统。该系统提供两级可选用户控制:参考音频可实现整体制作风格控制,可微分混音控制台则提供明确的音频效果参数以保障可解释性与细粒度优化。我们通过客观与主观评估,在混音质量与控制能力方面证明了系统的竞争力,项目页面提供代码与音频样本:this https URL。
英文摘要
Automatic music mixing aims to combine multitrack recordings into a balanced and coherent musical piece. Because the content of different songs and the subjective preferences of mixing engineers jointly shape the final outcome, a practical system should deliver well-balanced mixes while allowing for controllable stylistic variation. However, most existing methods treat automatic mixing and mixing style control as separate tasks, making it difficult for a single system to produce high-quality mixes while remaining editable and style-aware. To address this limitation, this paper presents Diff2Mix, a generative automatic mixing system based on diffusion models and a differentiable mixing console. This system offers two levels of optional user control: a reference audio enables overall production style control, and the differentiable mixing console provides explicit audio effects parameters for interpretability and fine-grained optimization. We demonstrate our system's competitive performance through both objective and subjective evaluations in terms of mixing quality and control ability. We provide code and audio samples at our project page https://zys711.github.io/Diff2Mix .
CommentsAccepted to ISMIR 2026