arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.02918cs.SDcs.LGeess.AS

学习爵士钢琴家风格:基于交叉注意力条件生成

Learning Jazz Pianist Style with Cross-Attention Conditioning

Drew Edwards, Akira Maezawa, Simon Dixon

中文总结 AI 辅助

本研究利用预训练符号音乐转换器编码钢琴家身份,通过交叉注意力条件生成特定风格音乐,并验证了风格捕捉的有效性及特征定位能力。

中文摘要 AI 辅助

爵士钢琴家会发展出独特的个人特征,有经验的听众往往能在几秒内识别出这些特征,然而支撑这种识别能力的底层特征却难以用形式化描述。我们通过预训练符号音乐转换器的视角研究爵士钢琴家风格,表明其学习到的表征已经足以编码钢琴家身份,并在两个基准上实现高精度分类。随后,我们在转换器中引入对学习到的钢琴家身份嵌入的交叉注意力,使其能够生成以特定艺术家风格为条件的音乐。两个评估协议证实了生成器捕捉到了有意义的风格结构:滑动窗口分类器持续将条件生成的续写归因于正确的艺术家,远高于无条件基线;而完全在合成生成上训练的分类器在12个类别中识别真实钢琴家,达到87%的片段级和95%的歌曲级准确率。最后,我们重新利用该分类器定位演奏中最具特征性的时刻,揭示区分每位钢琴家声音的具体音乐手势。

英文摘要

Jazz pianists develop distinctive traits that experienced listeners can often identify within seconds, yet the features underlying this recognition resist formal description. We study jazz pianist style through the lens of a pretrained symbolic music transformer, showing that its learned representations already encode pianist identity well enough for highly accurate classification across two benchmarks. We then augment the transformer with cross-attention over learned pianist identity embeddings, enabling it to generate music conditioned on a specific artist's style. Two evaluation protocols confirm that the generator captures meaningful stylistic structure: a sliding-window classifier consistently attributes conditioned continuations to the correct artist, far above unconditioned baselines; and a classifier trained entirely on synthetic generations identifies real pianists across 12 classes with 87% chunk-level and 95% song-level accuracy. Finally, we repurpose the classifier to locate the most characteristic moments within a performance, surfacing the specific musical gestures that distinguish each pianist's voice.

补充信息

↑