arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过跨模态自举学习音乐风格以实现钢琴编曲

Learning Music Style for Piano Arrangement Through Cross-Modal Bootstrapping

Jingwei Zhao, Gus Xia, Ziyu Wang, Ye Wang

arXiv 2608.03050首次发表:更新:

AI 中文总结

本文提出跨模态框架,受BLIP-2启发用Q-Former从预训练音频LM提取风格表示,经两阶段训练实现可控风格钢琴编曲,在多项任务中提升了风格对齐与音乐质量。

AI 中文摘要

什么是音乐风格?尽管常通过“摇摆乐”“古典”或“情感”等文本标签描述,但真正的风格隐含在具体音乐实例中。本文提出一种跨模态框架,从原始音频中学习隐含音乐风格并将其应用于符号音乐生成。受BLIP-2启发,该模型利用查询Transformer(Q-Former)从大型预训练音频语言模型(LM)中提取风格表示,进一步将其用于条件化符号LM以生成钢琴编曲。我们采用两阶段训练策略:对比学习对齐听觉风格与符号表达,随后进行生成式音乐编曲建模。我们的模型在主旋律谱(内容)和参考音频实例(风格)的共同条件下生成钢琴演奏,实现可控且风格忠实的编曲。实验验证了该方法在钢琴翻唱生成、风格迁移及音频转MIDI检索中的有效性,在风格感知对齐和音乐质量方面取得显著提升。

英文摘要

What is music style? Though often described using text labels such as "swing," "classical," or "emotional," the real style remains implicit and hidden in concrete music examples. In this paper, we introduce a cross-modal framework that learns implicit music styles from raw audio and applies them to symbolic music generation. Inspired by BLIP-2, our model leverages a Querying Transformer (Q-Former) to extract style representations from a large, pre-trained audio language model (LM), and further applies them to condition a symbolic LM for generating piano arrangements. We adopt a two-stage training strategy: contrastive learning to align auditory style with symbolic expression, followed by generative modeling for music arrangement. Our model generates piano performances jointly conditioned on a lead sheet (content) and a reference audio example (style), enabling controllable and stylistically faithful arrangement. Experiments demonstrate the effectiveness of our approach in piano cover generation, style transfer, and audio-to-MIDI retrieval, achieving substantial improvements in style-aware alignment and music quality.

CommentsAccepted by ISMIR 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑