另注:使用双流TFC-TDF U-Net与自适应集合归属的鲁棒性乐谱信息音符分离
On a Separate Note: Robust Score-Informed Note Separation with a Two-Stream TFC-TDF U-Net and Adaptive Set Ownership
查看机构详情
- Purdue University(普渡大学)
- Loyola University Chicago(芝加哥洛约拉大学)
- University of Michigan(密歇根大学)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本文提出首个基于深度学习的乐谱信息音符分离方法NoteSep,利用双流U-Net与自适应集合归属,在SCNS-Eval上中位SI-SDR达7.39dB,显著优于基线。
中文摘要 AI 辅助
乐谱信息音符分离旨在从多声道录音中提取所有单个音符的演奏波形。现有的深度学习系统通常仅针对乐器级音轨。我们提出了据我们所知首个用于乐谱信息音符分离的深度学习方法NoteSep。NoteSep通过应用提取阶段模型NoteGrab,对每个音符执行一次提取。NoteGrab以音高、起始和偏移为条件,通过双向交叉注意力连接的两个U-Net分离谐波和打击乐成分;选择性谐波门控抑制低八度干扰,同时保留打击乐攻击。最后,联合分离阶段应用自适应集合归属(ASO)来比较并发的NoteGrab估计并重新分配混合能量。我们构建了SCNS-Train(25,729个混合和743,920个目标)用于训练,以及SCNS-Eval(16种乐器,不重叠的乐谱和库)用于评估。在SCNS-Eval上,NoteSep的中位SI-SDR达到7.39分贝,而我们的最强基线为2.49分贝。演示页面见本https URL。
英文摘要
Score-informed note separation seeks to extract the performed waveform of all individual notes, often from a polyphonic recording. Existing deep learning systems generally only target instrument-level stems. We present, to our knowledge, the first deep learning approach to score-informed note separation, NoteSep. NoteSep extracts the queried notes by applying an extraction stage model, NoteGrab, once per note. Conditioned on pitch, onset, and offset, NoteGrab separates harmonic and percussive components in two U-Nets linked by bidirectional cross-attention; selective harmonic gating suppresses lower-octave interference while preserving percussive attacks. Finally, a joint separation stage applies Adaptive Set Ownership (ASO) to compare concurrent NoteGrab estimates and reallocate mixture energy. We curate SCNS-Train (25,729 mixtures and 743,920 targets) for training and SCNS-Eval (16 instruments, disjoint scores and libraries) for evaluation. On SCNS-Eval, NoteSep reaches a median SI-SDR of 7.39~dB, compared with 2.49~dB for our strongest baseline. See the demo page at https://benschou.com/notesep.