发表机构
Carnegie Mellon University(卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
介绍用于评估自动音乐转录模型的MulTTiPop数据集,通过对Lakh MIDI和TheoryTab数据集歌曲片段基于元数据匹配等方式收集,评估模型发现有改进空间,最佳模型起始F1得分为38%。
AI 中文摘要
我们展示了MulTTiPop,这是一个用于评估自动音乐转录模型的流行音乐片段及其相关多轨MIDI录音的数据集。MulTTiPop包含572个流行音乐片段,总计3.5小时音频,涵盖了从20世纪30年代到21世纪不同流派和年代的歌曲。为收集该数据集,我们对Lakh MIDI和TheoryTab数据集中的歌曲片段进行基于元数据的匹配,手动识别音频和MIDI之间的锚点节拍,然后对音频进行节拍跟踪并使MIDI变形以匹配其节奏和时间。我们在MulTTiPop上评估了最先进的自动音乐转录模型,发现仍有很大改进空间,最佳模型的起始F1得分为38%。更多MulTTiPop的详细信息和声音示例可在该https网址获取。
英文摘要
We present MulTTiPop, a dataset of pop music segments and their associated multitrack MIDI recordings for the evaluation of automatic music transcription models. MulTTiPop contains 572 segments of popular music totaling 3.5 hours of audio, and contains songs from diverse genres and decades from the 1930s to 2000s. To collect this dataset, we perform metadata-based matching on song segments from the Lakh MIDI and TheoryTab datasets, manually identify an anchor beat between the audio and MIDI, then use beat tracking on the audio and warp the MIDI to match its tempo and timing. We evaluate state-of-the-art automatic music transcription models on MulTTiPop and find substantial room for improvement, with the best model achieving 29% Onset F1. More details and sound examples of MulTTiPop are available at https://gclef-cmu.org/multtipop.
Comments8 pages, 4 figures. Associated web preview available at https://gclef-cmu.org/multtipop