PianoBind:面向流行钢琴音乐的多模态联合嵌入模型
PianoBind: A Multimodal Joint Embedding Model for Pop-piano Music
AI总结:
针对通用模型难捕捉同质性钢琴音乐细粒度语义、现有钢琴专用模型为单模态的问题,提出钢琴专用多模态联合嵌入模型PianoBind,优化后在文本到音乐检索任务上表现优于通用模型,设计思路可复用。
AI中文摘要:
独奏钢琴音乐虽为单乐器媒介,却具备强大的表现力,可传递跨流派、情绪与风格的丰富语义信息。然而,当前通用音乐表示模型主要基于大规模数据集训练,往往难以捕捉同质性独奏钢琴音乐内部细微的语义差异。此外,现有钢琴专用表示模型通常为单模态,无法捕捉钢琴音乐通过音频、符号和文本模态所呈现的固有多模态特性。为解决这些局限,我们提出PianoBind,一种钢琴专用的多模态联合嵌入模型。我们在联合嵌入框架内系统研究了多源训练与模态利用策略,该框架针对捕捉(1)小规模及(2)同质性钢琴数据集中的细粒度语义差异进行了优化。实验结果表明,PianoBind学习到的多模态表示能有效捕捉钢琴音乐的细微差别,在域内和域外钢琴数据集上的文本到音乐检索性能均优于通用音乐联合嵌入模型。此外,我们的设计选择为钢琴音乐之外的同质性数据集多模态表示学习提供了可复用的思路。
英文摘要:
Solo piano music, despite being a single-instrument medium, possesses significant expressive capabilities, conveying rich semantic information across genres, moods, and styles. However, current general-purpose music representation models, predominantly trained on large-scale datasets, often struggle to captures subtle semantic distinctions within homogeneous solo piano music. Furthermore, existing piano-specific representation models are typically unimodal, failing to capture the inherently multimodal nature of piano music, expressed through audio, symbolic, and textual modalities. To address these limitations, we propose PianoBind, a piano-specific multimodal joint embedding model. We systematically investigate strategies for multi-source training and modality utilization within a joint embedding framework optimized for capturing fine-grained semantic distinctions in (1) small-scale and (2) homogeneous piano datasets. Our experimental results demonstrate that PianoBind learns multimodal representations that effectively capture subtle nuances of piano music, achieving superior text-to-music retrieval performance on in-domain and out-of-domain piano datasets compared to general-purpose music joint embedding models. Moreover, our design choices offer reusable insights for multimodal representation learning with homogeneous datasets beyond piano music.