arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.01093cs.SD

分离与检测:通过潜在扩散实现统一鼓转录与音轨生成

Separate-and-Detect: Unified Drum Transcription and Stem Generation via Latent Diffusion

Wei-Han Hsu, Chih-Cheng Chang, Bo-Yu Chen, Li Su, Yi-Hsuan Yang

AI总结:

本文提出分离与检测框架,以5音轨潜在扩散模型为前端,结合起始点分支等辅助训练,实现鼓转录与音轨生成,性能优于基线,兼具分离音轨能力。

AI中文摘要:

自动鼓转录(ADT)通常被建模为从音乐混合音频到符号鼓事件的直接映射,该方法虽能有效完成转录,但会丢失对编辑、混音及制作有用的声学音轨。本文重新探讨替代的分离与检测框架:鼓源分离前端先生成5种可编辑鼓音轨,固定起始点检测器再将各音轨转换为符号事件。该分离器基于5音轨潜在扩散模型构建,在紧凑的VAE潜在空间中联合生成底鼓、军鼓、嗵鼓、踩镲和镲片。本文进一步研究两个仅用于训练的辅助分支——起始点分支(OB)和音色分支(TB),二者在学习阶段优化分离器,推理阶段则被舍弃。在合成鼓多轨音频上训练,在MDB Drums和ENST-Drums数据集上评估,所提流程在整体转录F1值上始终优于基于U-Net的强鼓分离基线;在评估协议下,其底鼓和军鼓F1值也优于代表性端到端ADT系统,同时还能提供分离的音频音轨。 ablation结果显示,OB带来最稳定的转录增益,TB则改变重建、声学音轨质量与起始点检测间的权衡。这些结果表明,生成式鼓分离不仅可作为源分离模型,还可作为可解释鼓转录的实用前端。

英文摘要:

Automatic Drum Transcription (ADT) is commonly formulated as a direct mapping from a music mixture to symbolic drum events. While effective for transcription, this formulation discards the acoustic stems that are useful for editing, remixing, and production. We revisit an alternative separate-and-detect formulation, where a drum source separation front end first produces five editable drum stems, and a fixed onset detector then converts each stem into symbolic events. The separator is built on a five-stem latent diffusion model that jointly generates kick, snare, toms, hi-hats, and cymbals in a compact VAE latent space. We further study two training-only auxiliary branches--an onset branch (OB) and a timbre branch (TB)--which shape the separator during learning but are discarded at inference. Trained on synthetic drum multitracks and evaluated on MDB Drums and ENST-Drums, the proposed pipeline consistently improves over a strong U-Net-based drum separation baseline in overall transcription F1. It also outperforms a representative end-to-end ADT system on kick and snare F1 under our evaluation protocol, while additionally providing separated audio stems. The ablation results show that OB gives the most stable transcription gains, whereas TB changes the trade-off between reconstruction, acoustic stem quality, and onset detection. These results suggest that generative drum demixing can serve not only as a source separation model, but also as a practical front end for interpretable drum transcription.

补充信息

↑