发表机构
National Technical University of Athens; Athena Research Center(雅典国家技术大学; 雅典娜研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对希腊传统音乐自动标注,提出利用舞者姿态作为音频之外的补充信息,通过多模态融合(音频、视频、骨架)提升性能,最佳三模态系统较音频基线宏ROC-AUC提升约4个百分点。
AI 中文摘要
自动标注是音乐信息检索(MIR)中的核心任务,然而大多数标注系统仅利用音频。现场音乐表演本质上是多模态的,因为诸如乐器、地域风格和舞蹈形式等语义标签同时编码在声学、视觉和具身表演线索中。对于文化特定的曲目,如希腊传统音乐,尤其如此,这些曲目在MIR基准中代表性不足。在本文中,我们研究了舞者姿态是否能为希腊传统音乐的自动标注提供音频之外的补充信息。使用Lyra数据集,我们通过提取对齐的视频特征和姿态衍生的骨架流来扩展先前仅基于音频的工作,从而实现了多模态自动标注的实验设置。我们进一步引入了一个自动化流程,用于从野外舞蹈视频中提取主舞者骨架序列,该流程结合了舞蹈场景检测、多人跟踪、舞者选择、姿态估计和质量过滤。我们使用多种融合策略比较了单模态、所有双模态组合和三模态系统。音频仍然是最强的单一模态(AST:宏ROC-AUC 0.821),而骨架虽然单独使用时较弱,但通过多模态融合提升了性能。最佳三模态系统相比最强的音频基线,宏ROC-AUC提高了约4个百分点。
英文摘要
Automatic tagging is a core task in Music Information Retrieval (MIR), yet most tagging systems exploit only audio. Live music performance is inherently multimodal, as semantic labels such as instruments, regional styles, and dance forms are encoded simultaneously across acoustic, visual, and embodied performance cues. This is especially true of culturally specific repertoires such as Greek traditional music, which remain underrepresented in MIR benchmarks. In this paper, we investigate whether the use of dancer pose provides complementary information for automatic tagging in Greek traditional music beyond audio. Using the Lyra dataset, we extend prior audio-only work by extracting aligned video features and pose-derived skeleton streams, enabling an experimental setting for multimodal auto-tagging. We further introduce an automated pipeline for extracting primary-dancer skeleton sequences from in-the-wild dance footage, combining dance-scene detection, multi-person tracking, dancer selection, pose estimation, and quality filtering. We compare unimodal, all bimodal combinations, and trimodal systems using multiple fusion strategies. Audio remains the strongest single modality (AST: macro ROC-AUC 0.821), while skeletons, though weak in isolation, enhance performance through multimodal fusion. The best trimodal system improves macro ROC-AUC by about 4 percentage points over the strongest audio baseline.