SPHERE:通过结合空间启发式奖励的音频语言模型后训练实现自动音乐上混
SPHERE: Automatic Music Upmixing via Audio Language Model Post-Training with Spatial Heuristic Rewards
浏览论文内容
中文总结 AI 辅助
该研究提出基于音频语言模型后训练的自动音乐上混方法,设计含6个子奖励的Sphere奖励套件,证明领域知识可蒸馏入语言模型实现该任务。
中文摘要 AI 辅助
本文研究自动音乐上混任务,即系统从多轨录音预测空间混音参数。与依赖特定任务音乐编码器的现有方法不同,我们通过音频语言模型(ALM)后训练处理该任务,利用现有ALM编码音乐语义和混音知识的丰富表示。具体而言,我们提出一种后训练方案,先采用拒绝采样监督微调(SFT),再通过带可验证奖励的强化学习(RLVR)结合GRPO进行训练。我们提出Sphere(空间启发式奖励),一套受音乐混音惯例启发的确定性奖励套件,用于指导后训练,它包含6个感知驱动的子奖励,鼓励输出混音居中、平衡且有空间感。更广泛而言,我们的结果表明,专业领域知识可编码为可验证奖励并蒸馏到语言模型中,无需特定任务架构。
英文摘要
In this paper, we study the task of automatic music upmixing, wherein a system predicts spatial mixing parameters from a multi-stem recording. Different from existing methods that rely on task-specific music encoders, we approach this task via audio language model (ALM) post-training, leveraging rich representations from existing ALMs, which encode both music semantics and mixing knowledge. Specifically, we propose a post-training recipe that first employs rejection sampling SFT, followed by reinforcement learning (RL) with verifiable rewards (RLVR) via GRPO. We propose Sphere (Spatial Heuristic Rewards), a deterministic reward suite inspired by music mixing conventions, to guide our post-training. It consists of 6 perceptually-motivated sub-rewards and encourages the output mix to be centered, balanced and spacious. More broadly, our results suggest that expert domain knowledge can be encoded as verifiable rewards and distilled into language models, without task-specific architectures.
发表机构
- Meta Reality Labs(元宇宙实验室)
- Queen Mary University of London(伦敦玛丽女王大学)
机构由 AI 辅助整理,请以论文原文为准。