发表机构
The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
CPR框架结合连续自回归建模、局部流匹配和上采样精修,实现高质量钢琴MIDI渲染,并引入BREPA和MT-RoPE增强语义与对齐。
AI 中文摘要
提示条件钢琴MIDI到音乐渲染旨在忠实渲染目标音符,同时再现参考录音的音色。现有方法主要遵循两种范式:自回归(AR)建模和流匹配(或扩散)。离散码本自回归模型提供因果时间建模,但量化可能丢弃声学细节。流匹配在牺牲全序列注意力成本和更差语义结构的情况下,更好地保留声学结构。连续自回归模型直接操作连续表示。它不仅结合了AR模型的条件跟随能力和流匹配的分布建模能力,还绕过了量化瓶颈,计算成本更低。基于这一原理,我们提出了作曲家-演奏者-精修者(CPR)框架。作曲家自回归预测连续隐藏状态,演奏者通过局部流匹配生成24kHz声学潜变量,精修者随后将波形上采样至48kHz。我们进一步引入瓶颈表示对齐(BREPA)和模态-时间旋转位置编码(MT-RoPE),以增强作曲家隐藏状态中的音乐语义结构和跨模态的时间对齐。代码可在https://this URL获取。
英文摘要
Prompt-conditioned piano MIDI-to-Music rendering aims to faithfully render target notes while reproducing the timbre of a reference recording. Existing approaches primarily follow two paradigms: autoregressive (AR) modeling and flow matching (or diffusion). Discrete-codec AR models provide causal temporal modeling, but quantization can discard acoustic detail. Flow matching better preserves acoustic structure in the cost of full-sequence attention costs and worse semantic structure. Continuous autoregressive models operate directly on continuous representations. It not only combines the condition-following ability of AR models and distribution-modeling capacity of flow matching but also bypasses the quantization bottleneck with lower computational costs. Building on this principle, we present Composer--Performer--Refiner (CPR) framework. Composer autoregressively predicts continuous hidden states, Performer generates 24kHz acoustic latents through local flow matching and Refiner then upsamples the waveform to 48 kHz. We further introduce Bottlenecked Representation Alignment (BREPA) and Modality--Time RoPE (MT-RoPE) to strengthen musical semantic structure in Composer hidden states and temporal alignments across modalities. Codes are available at https://github.com/FEAfeatherTHER/CPR_official