arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.06488cs.SDeess.AS

基于音素对齐的声乐合奏混合中主唱分离

Lead Vocal Separation from Vocal Ensemble Mixtures Using Phoneme Alignment

  • The University of Tokyo(东京大学)
  • National Institute of Advanced Industrial Science and Technology (AIST)(独立行政法人产业技术综合研究所(AIST))

机构由 AI 辅助整理,请以论文原文为准。

Yuma Narahata, Tomohiko Nakamura, Yuki Saito, Hiroshi Saruwatari

AI总结:

针对声乐合奏中主唱分离难的问题,提出基于BS-RoFormer和音素对齐条件化的模型,利用FiLM注入帧级音素标签,显著提升分离性能。

AI中文摘要:

当代无伴奏合唱常呈现主唱与伴奏的织体结构,其中主唱声部(Vo)承载主旋律,其余声部提供伴奏。由于角色不同,将主唱声部从其余声部中分离出来(称为Vo分离)可实现歌词识别和无伴奏合唱音乐的伴奏生成等下游应用。尽管有这些潜在应用,该任务的声学线索有限,因为目标源和干扰源均为声学特征相似且常时间重叠的歌唱人声,使Vo分离具有挑战性。本文提出一种利用Vo声部音素对齐作为辅助信息的Vo分离模型。该模型基于带分裂RoPE Transformer(BS-RoFormer)——一种最先进的音乐源分离模型,并通过特征级线性调制(FiLM)将帧级音素标签引入其中间表示。实验结果表明,音素对齐条件化相比仅音频基线提升了Vo分离性能,且比仅以Vo歌唱/静音活动为条件获得更大的平均增益。进一步分析表明,当剩余声部中与Vo共享相同音素的声部较少时,音素标签信息的优势更大。

英文摘要:

Contemporary a cappella singing often has a lead-and-accompaniment texture, where the lead vocal (Vo) part carries the main melody and the remaining vocal parts provide accompaniment. Owing to their distinct roles, separating the Vo part from the remaining vocal parts, referred to as Vo separation, enables downstream applications such as lyric recognition and minus-one accompaniment generation for vocal ensemble music. Despite these potential applications, acoustic cues for this task are limited because the target and interfering sources are all singing voices with similar acoustic characteristics and often overlap in time, making Vo separation challenging. In this paper, we propose a Vo separation model that uses phoneme alignment of the Vo part as auxiliary information. The proposed model is based on band-split RoPE Transformer (BS-RoFormer), a state-of-the-art music source separation model, and introduces frame-level phoneme labels into its intermediate representations using feature-wise linear modulation (FiLM). Experimental results show that phoneme-alignment conditioning improves Vo separation performance over an audio-only baseline and yields larger average gains than conditioning only on Vo singing/silence activity. Further analysis suggests that the advantage of phoneme-label information is larger when fewer remaining vocal parts share the same phoneme as Vo.

补充信息

↑