PianoVAM:多模态钢琴演奏数据集
PianoVAM: A Multimodal Piano Performance Dataset
AI总结:
本文提出PianoVAM多模态钢琴演奏数据集,包含视频、音频、MIDI、手部关键点及指法标签,并给出音视频钢琴转录基准结果。
AI中文摘要:
音乐演奏的多模态特性,推动了音乐信息检索(MIR)社区对音频领域之外数据的日益增长的兴趣。本文介绍了PianoVAM,一个全面的钢琴演奏数据集,包含视频、音频、MIDI、手部关键点、指法标签和丰富的元数据。该数据集使用Disklavier钢琴录制,在真实且多样的演奏条件下,捕捉业余钢琴家日常练习期间的音频和MIDI,同时同步录制俯视视频。手部关键点和指法标签通过预训练的手部姿态估计模型和半自动指法标注算法提取。我们讨论了数据收集过程中遇到的挑战以及不同模态之间的对齐过程。此外,我们描述了基于从视频中提取的手部关键点的指法标注方法。最后,我们展示了使用PianoVAM数据集进行仅音频和音视频钢琴转录的基准测试结果,并讨论了其他潜在应用。
英文摘要:
The multimodal nature of music performance has driven increasing interest in data beyond the audio domain within the music information retrieval (MIR) community. This paper introduces PianoVAM, a comprehensive piano performance dataset that includes videos, audio, MIDI, hand landmarks, fingering labels, and rich metadata. The dataset was recorded using a Disklavier piano, capturing audio and MIDI from amateur pianists during their daily practice sessions, alongside synchronized top-view videos in realistic and varied performance conditions. Hand landmarks and fingering labels were extracted using a pretrained hand pose estimation model and a semi-automated fingering annotation algorithm. We discuss the challenges encountered during data collection and the alignment process across different modalities. Additionally, we describe our fingering annotation method based on hand landmarks extracted from videos. Finally, we present benchmarking results for both audio-only and audio-visual piano transcription using the PianoVAM dataset and discuss additional potential applications.