MeloCodec:利用旋律先验实现高保真歌声表征
MeloCodec: Harnessing Melodic Priors for High-Fidelity Singing Voice Representation
浏览论文内容
中文总结 AI 辅助
MeloCodec是整合旋律先验的新型音频编解码器框架,采用先分词后融合范式与两阶段训练策略,在歌声表征任务中优于基线,提升了音高一致性并实现可控音高操作。
中文摘要 AI 辅助
神经音频编解码器是基于大语言模型(LLM)的音频生成的基础分词器。尽管语义先验被广泛用于提升语言可理解性,但明确声学先验的整合仍未得到充分探索,这限制了频率敏感领域的合成保真度。为解决这一差距,我们提出MeloCodec,这是一种旨在有效整合旋律先验的新框架,旋律先验是歌声的关键声学信息形式。为解决此类明确先验直接融合通常导致的优化不稳定性,我们提出“先分词后融合”范式,即预训练离散旋律分支以锁定结构后再进行特征融合。为稳健实现该范式,我们进一步提出两阶段训练策略,以防止码本崩溃并确保稳定收敛。实验表明,MeloCodec在歌声表征方面优于基线模型,提升了音高一致性,并能在最小音色退化的情况下实现可控音高操作。
英文摘要
Neural audio codecs serve as fundamental tokenizers for LLM-based audio generation. While semantic priors are widely exploited to enhance linguistic intelligibility, the integration of explicit acoustic priors remains underexplored, limiting synthesis fidelity in frequency-sensitive domains. To address this gap, we introduce MeloCodec, a novel framework designed to effectively incorporate melodic priors, a critical form of acoustic information for singing. To address the optimization instability typically caused by the direct fusion of such explicit priors, we propose a Tokenize-then-Fuse paradigm that pre-trains a discrete melodic branch to lock in structures before feature fusion. To robustly realize this paradigm, we further propose a two-stage training strategy that prevents codebook collapse and ensures stable convergence. Experiments show that MeloCodec outperforms baselines in singing voice representation, improving pitch consistency and enabling controllable pitch manipulation with minimal timbre degradation.