arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向声部感知的合唱转录与歌声声部分配

Toward Part-Aware Choral Transcription with singing voice assignment

Hanyu Meng, Zhanhong He, Zixun Guo, Yaolong Ju

arXiv 2610.09861首次发表:更新:

发表机构

The University of New South Wales; The University of Western Australia; Queen Mary University of London; Great Bay University(新南威尔士大学; 西澳大利亚大学; 伦敦玛丽女王大学; 大湾区大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出首个端到端声部感知合唱转录框架PawCT,联合建模音符转录与SATB声部分配,在YouChorale上显著优于现有基线,验证了联合建模的有效性。

AI 中文摘要

单乐器自动音乐转录(AMT)已取得显著进展,但合唱应用要求将女高音、女低音、男高音和男低音(SATB)转录为独立声部。现有的音符级合唱AMT方法仅生成单一的合并音符轨道,限制了排练、教学和乐谱重建。为解决此局限,我们提出了面向声部感知的合唱转录(PawCT),据我们所知,这是首个从合唱音频中识别活跃SATB声部并将其分别转录为独立音符级轨道的端到端神经框架。PawCT结合了声部特定的起音、偏移和帧预测,以及声部存在性估计、联合级监督和基于SATB音域的先验(RP)及其有序连续性(OC)扩展的结构化训练目标,后者增加了声部内旋律连续性和跨声部音高排序。在YouChorale上,PawCT-RP-OC在50毫秒起音容差下实现了0.225的宏观声部感知音符F1值,相对优于改编的合唱基线(0.165)36.4%,并优于两阶段后处理分配流程(0.175)。其声部无关变体PagCT实现了0.382的50毫秒起音F1值,而先前最先进的合唱AMT模型为0.237。在CSD和Cantoria上的跨数据集评估进一步检验了数据集偏移下的性能。这些结果证明了联合建模音符转录和声部分配的优势。代码和演示可在该https URL获取。

英文摘要

Single-instrument automatic music transcription (AMT) has advanced substantially, yet choral applications require soprano, alto, tenor, and bass (SATB) to be transcribed as separate parts. Recent note-level choral AMT instead produces a single merged note track, limiting rehearsal, education, and score reconstruction. To address this limitation, we introduce Part-aware Choral Transcription (PawCT), to our knowledge the first end-to-end neural framework that identifies active SATB parts from choral audio and transcribes each into a separate note-level track. PawCT combines part-specific onset, offset, and frame prediction with part-presence estimation, union-level supervision, and structured training targets using a range prior (RP) based on SATB pitch ranges and its ordered-continuity (OC) extension, which adds within-part melodic continuity and cross-part pitch ordering. On YouChorale, PawCT-RP-OC achieves a macro part-aware note F1 of 0.225 at a 50-ms onset tolerance, outperforming an adapted choral baseline (0.165) by 36.4% relative and a two-stage post-hoc assignment pipeline (0.175). Its part-agnostic variant, PagCT, achieves a 50-ms onset F1 of 0.382, compared with 0.237 for the previous state-of-the-art choral AMT model. Cross-dataset evaluations on CSD and Cantoria further assess performance under dataset shift. These results demonstrate the benefit of jointly modeling note transcription and vocal-part assignment. Code and demos are available at https://hanyu-meng.github.io/Paw_Choral_AMT_Demo/.

CommentsSubmitted to ICASSP2027

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑