SCISSOR:面向管弦乐录音的乐谱条件乐器源分离
SCISSOR: Score-Conditioned Instrument Source Separation for Orchestral Recordings
浏览论文内容
中文总结 AI 辅助
SCISSOR利用乐谱生成逐帧查询,通过共享音频表示和softmax分配时频证据,实现管弦乐录音的乐器源分离,在真实录音上取得最高SDR,且对乐谱损坏鲁棒。
中文摘要 AI 辅助
管弦乐分离从混合信号中恢复乐器声部,其中共享的音高、谐波和音色模糊了源的身份。对齐的乐谱提供了乐器标签、音符音高和活动时间。一种基于乐谱的方法在掩码预测之前将钢琴卷附加到音频特征上。我们提出了SCISSOR(面向管弦乐录音的乐谱条件乐器源分离),它利用乐谱为每个源形成逐帧查询。每个查询匹配一个共享的音频表示,并且对乐器和背景槽位的softmax联合分配重叠的时频证据。即使乐谱中缺少音符,查询也能保留乐器身份。在SynthSOD和一小部分URMP及PHENICX-Anechoic录音上训练后,SCISSOR在保留的真实录音上达到了最高的平均SDR。仅使用SynthSOD训练时,它在SynthSOD和零样本PHENICX-Anechoic上领先,并在零样本URMP上优于其仅音频的对照组。在乐谱损坏情况下,SCISSOR的退化程度也小于所评估的基于乐谱的基线方法。
英文摘要
Orchestral separation recovers instrument sections from mixtures in which shared pitches, harmonics, and timbres obscure source identity. An aligned score provides instrument labels, note pitches, and activity times. A score-informed approach appends piano rolls to audio features before mask prediction. We introduce SCISSOR (Score-Conditioned Instrument Source Separation for Orchestral Recordings), which uses the score to form a frame-wise query for each source. Each query matches a shared audio representation, and a softmax over instrument and background slots jointly assigns overlapping time-frequency evidence. The queries retain instrument identity even when notes are missing from the score. After training on SynthSOD and a small set of URMP and PHENICX-Anechoic recordings, SCISSOR achieves the highest average SDR on held-out real recordings. With SynthSOD-only training, it leads on SynthSOD and zero-shot PHENICX-Anechoic, and improves on its audio-only control on zero-shot URMP. SCISSOR also degrades less under score corruption than the evaluated score-based baselines.