发表机构
University of Illinois Urbana-Champaign; Siebel School of Computing and Data Science(伊利诺伊大学厄巴纳-香槟分校; 西贝尔计算与数据科学学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对多歌手音乐混合信号,提出结合目标歌手注册录音嵌入的人声源分离框架,在二重唱数据集上验证了其可显著提升目标歌手提取的SI-SDR与感知质量。
AI 中文摘要
音乐源分离系统通常仅提取单一人声轨道,无法区分多位歌手。本研究针对多歌手混合信号开展基于歌手信息的人声源分离研究,所提出的框架引入目标歌手的短时长注册录音,通过学习得到的嵌入向量引导分离过程;该歌手嵌入向量通过特征拼接或特征-wise线性调制(FiLM)方式融入模型,使模型能够聚焦于目标歌手,同时抑制干扰。研究人员基于DAMP-VSEP构建二重唱数据集,开展质量过滤并设置非重叠的注册片段。在独唱与二重唱场景下的实验表明,基线模型在单歌手混合信号上表现良好,而所提方法可提升多歌手场景下的目标歌手提取效果,将目标歌手的SI-SDR从0.33 dB提升至5.58 dB;Fréchet音频距离(FAD)的结果进一步表明,该方法可提升感知质量,且与目标音频分布的对齐度更好。代码与检查点可在指定URL获取。
英文摘要
Music source separation systems typically extract a single vocal track and do not distinguish between multiple singers. We study singer-informed vocal source separation for multi-singer mixtures. Our framework introduces a short enrollment recording of a target singer to guide separation through a learned embedding. The singer embedding is incorporated using feature concatenation or feature-wise linear modulation (FiLM), enabling the model to focus on the target singer while suppressing interference. We construct a duet dataset based on DAMP-VSEP with quality filtering and non-overlapping enrollment segments. Experiments on solo and duet settings show that while baseline models perform well for single-singer mixtures, the proposed method improves target-singer extraction in multi-singer cases, increasing target-singer SI-SDR from 0.33 dB to 5.58 dB. Fréchet Audio Distance (FAD) further shows improved perceptual quality and better alignment with target audio distributions. Code and checkpoints are available at https://github.com/jocelynxu01/singer-separation-paper.
CommentsAccepted at IWAENC 2026