arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.19776cs.SDcs.AI

RPPNet:通过边界感知建模生成长期结构旋律的感知分组节奏-音高基元

RPPNet: Perceptually-Grouped Rhythm-Pitch Primitives for Long-Term Structure Melody Generation via Boundary-Aware Modeling

Tieyao Zhang, Yuke Liu, Jiaxing Yu, Xinda Wu, Kejun Zhang, Genfang Chen

首次发表
浏览论文内容

中文总结 AI 辅助

本文针对现有符号音乐生成模型的问题,提出RPPNet架构,通过生成可变长度的RPP序列并解码,其分组基于音乐心理学自动得出。实验显示该模型生成的旋律在长期结构和音乐性上更优,消融研究表明性能提升源于心理表征结构正确,为音乐生成提供跨学科视角。

中文摘要 AI 辅助

现有符号音乐生成模型通常以小节为基本结构单元,但人类对乐句的感知常与标注的小节线不一致,导致长期结构碎片化。本文提出RPPNet,一种具有可变结构边界的两阶段深度学习架构。首先生成可变长度的节奏-音高基元(RPP)序列,每个RPP编码音符数量、节奏和轮廓,然后将RPP序列解码为具体音符。RPP的分组基于音乐心理学从声学线索、听觉惯性和相似性感知自动得出。实验表明,RPPNet生成的旋律在长期结构和音乐性方面都更优,在所有主观评价维度上都有显著改善。消融研究证实性能提升源于心理表征的结构正确性,而非模型容量。这项工作为音乐生成提供了跨学科视角,整合了音乐理论、计算建模和音乐心理学。

英文摘要

Existing symbolic music generation models typically use bars as the basic structural unit. However, human perception of musical phrases often does not align with notated bar lines, leading to long-term structural fragmentation. This paper proposes RPPNet-a two-stage deep learning architecture with variable structural boundaries. It first generates variable-length Rhythm-Pitch Primitive (RPP) sequences, where each RPP encodes note count, rhythm, and contour; then decodes the RPP sequences into concrete notes. The grouping of RPPs is automatically derived from acoustic cues, auditory inertia, and similarity perception based on music psychology. Experiments show that melodies generated by RPPNet are superior in both long-term structure and musicality, with significant improvements across all subjective evaluation dimensions. Ablation studies confirm that the performance gain stems from the structural correctness of the psychological representation, rather than from model capacity. This work offers an interdisciplinary perspective for music generation, integrating music theory, computational modeling, and music psychology.

↑