arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2601.21740cs.MMcs.SD

MIDI-LLaMA:一种用于符号音乐理解的指令遵循多模态大语言模型

MIDI-LLaMA: An Instruction-Following Multimodal LLM for Symbolic Music Understanding

  • SensiLab, Monash University, Australia(Monash大学)
  • University of Sussex, Brighton, United Kingdom(Sussex大学)
  • School of Computing and Information Systems, The University of Melbourne, Australia(墨尔本大学计算机与信息系统学院)

机构由 AI 辅助整理,请以论文原文为准。

Meng Yang, Jon McCormack, Maria Teresa Llano, Wanchao Su, Chao Lei

更新

AI总结:

MIDI-LLaMA通过结合MusicBERT和Llama-3-8B,实现了对符号音乐的指令遵循理解,显著提升了音乐描述和语义对齐能力。

AI中文摘要:

近期在音频音乐的多模态大语言模型(MLLM)方面取得了进展,展示了在音乐理解方面的强大能力,但符号音乐,作为音乐结构的基本表示,仍然未被探索。在本工作中,我们介绍了MIDI-LLaMA,这是首个用于符号音乐理解的指令遵循MLLM。我们的方法通过包含特征对齐和指令微调的两阶段流程,将MIDI编码器MusicBERT与Llama-3-8B对齐。为了支持训练,我们设计了一个可扩展的标注流程,对GiantMIDI-Piano进行细粒度元数据标注,生成了一个MIDI-文本数据集。与在相同指令微调过程中训练于将MIDI转换为ABC记谱法的基线模型相比,MIDI-LLaMA在描述和问答中的语义对齐方面表现显著优于基线。人类评估进一步确认了MIDI-LLaMA在音乐理解、情绪识别、创造力和整体偏好方面的优势。这些发现表明,将符号音乐纳入大语言模型中增强了其音乐理解能力。

英文摘要:

Recent advances in multimodal large language models (MLLM) for audio music have demonstrated strong capabilities in music understanding, yet symbolic music, a fundamental representation of musical structure, remains unexplored. In this work, we introduce MIDI-LLaMA, the first instruction-following MLLM for symbolic music understanding. Our approach aligns the MIDI encoder MusicBERT and Llama-3-8B via a two-stage pipeline comprising feature alignment and instruction tuning. To support training, we design a scalable annotation pipeline that annotates GiantMIDI-Piano with fine-grained metadata, resulting in a MIDI-text dataset. Compared with the baseline trained on converting MIDI into ABC notation under the same instruction-tuning procedure, MIDI-LLaMA substantially outperforms in captioning and semantic alignment in question answering. Human evaluation further confirms the advantages of MIDI-LLaMA in music understanding, emotion recognition, creativity, and overall preference. These findings demonstrate that incorporating symbolic music into large language models enhances their capacity for musical understanding.

补充信息

↑