AI 中文总结
研究如何利用全局监督实现局部多模态音乐对齐,提出FuSiLi方法,通过基于Sinkhorn的软对齐操作局部特征,结合传统全局相似性微调预训练编码器,在跨模态检索和帧级对齐任务中表现出色,优于多种基线。
AI 中文摘要
理解音乐需要理解跨数据模态的局部关系,例如演奏音频中的时间如何映射到乐谱图像中的位置。然而,获得这种局部对应关系的监督很困难,实际中通常只有更粗略的全局监督,如音频和图像的配对片段。为解决这一差距,我们提出FuSiLi(融合Sinkhorn-局部相似性),一种用于多模态对比学习的相似性分数,通过基于Sinkhorn的软对齐直接对局部图像块和音频帧特征进行操作。我们证明FuSiLi(i)有效学习局部关系,(ii)仅需全局监督,(iii)保留传统对比方法的全局对齐能力。我们使用结合FuSiLi与传统全局相似性的混合对比目标,对成对的原始乐谱图像和音频微调预训练的CLIP和CLAP编码器。我们在跨模态检索和帧级对齐任务中针对一系列全局和局部基线进行评估,表明我们的方法在局部对齐上优于它们,在检索上保持竞争力。
英文摘要
Understanding music requires understanding localized relationships across data modalities, e.g., how time in performance audio maps onto position in a score image. Yet supervision for such local correspondences is difficult to obtain-in practice, we often only have access to coarser global supervision like paired segments of audio and images. To address this gap, we propose FuSiLi (Fused Sinkhorn-Localized Similarity), a similarity score for multimodal contrastive learning operating directly on local image patch and audio frame features via Sinkhorn-based soft alignment. We show that FuSiLi (i) effectively learns local relationships, (ii) requires only global supervision, and (iii) retains the global alignment capabilities of conventional contrastive approaches. We fine-tune pretrained CLIP and CLAP encoders on pairs of raw sheet music images and audio using a hybrid contrastive objective combining FuSiLi with conventional global similarity. We evaluate on cross-modal retrieval and frame-level alignment tasks against a range of global and local baselines, showing that our approach outperforms them on local alignment while remaining competitive on retrieval.
CommentsISMIR 2026