发表机构
Sogang University; KAIST; NYU Shanghai; JIUTIAN Research, China Mobile; The State Key Laboratory of Multimedia Information Processing, Peking University; China Mobile (Hong Kong) Innovation Research Institute; Central Conservatory of Music(西江大学; 韩国科学技术院; 上海纽约大学; 中国移动九天研究院; 北京大学多媒体信息处理国家重点实验室; 中国移动(香港)创新研究院; 中央音乐学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对低资源音乐理解问题,推出UniVerse基准与数据集,训练LALMs并研究多模态不平衡学习策略,实现性能提升但仍存在深层音乐理解的差距。
AI 中文摘要
近期大型音频-语言模型(Large Audio-Language Models, LALMs)在音乐字幕生成、流派分类、声音事件检测等任务上的性能已显著提升,但针对其在不同音乐传统(尤其是植根于独特文化语境的民间音乐)间适应性的改进却关注不足。民间音乐传统通常资源匮乏、区域分布不均且记录不足;即便此类样本出现在大规模预训练中,LALMs也常无法捕捉其结构与风格特征,部分原因在于缺乏专用评估协议与训练方案。为解决这些局限,我们推出UniVerse——一种用于低资源音乐理解的可复现解决方案。具体而言,我们构建了UniVerseBench:一个包含5042个问答对、覆盖38个以上文化与语言实体的基准,通过专家指导但高度自动化的流程生成;同时构建了完全自动化、由模型生成的多轮对话训练数据集UniVerseSet。通过在UniVerseSet上训练LALMs,我们系统地在稠密架构与混合专家(Mixture-of-Experts, MoE)架构上适配并研究了代表性的多模态不平衡学习策略。实验结果表明,完全自动化的数据整理结合感知不平衡的训练可带来显著提升,但模型仍难以捕捉细粒度声学特征,这表明表层对齐与深层音乐理解之间存在差距。
英文摘要
Recent advances in large audio-language models (LALMs) have significantly improved performance in tasks such as music captioning, genre classification, and sound event detection. However, limited attention has been paid to improving their adaptability across diverse musical traditions, particularly folk music rooted in distinct cultural contexts. Folk-music traditions are typically resource-scarce, unevenly represented across regions, and poorly documented. Even when such samples appear in large-scale pre-training, LALMs often fail to capture their structural and stylistic characteristics, partly due to the absence of dedicated evaluation protocols and training solutions. To address these limitations, we introduce UniVerse, a reproducible solution for low-resource music understanding. Specifically, we propose UniVerseBench, a benchmark of 5,042 Q&A pairs across more than 38 cultural and linguistic entities, constructed via an expert-guided yet highly automated pipeline. In parallel, we construct a fully automated, model-generated multi-turn dialogue training dataset UniVerseSet. By training LALMs on UniVerseSet, we systematically adapt and investigate representative multimodal imbalance learning strategies across both dense and Mixture-of-Experts (MoE) architectures. Experimental results indicate that fully automated data curation combined with imbalance-aware training yields non-trivial improvements, but models still struggle to capture fine-grained acoustic features, indicating a gap between surface-level alignment and deep musical comprehension.
Comments21 pages, 7 figures, 8 tables