发表机构
Northwestern Polytechnical University; Geely Automobile Research Institute (Ningbo) Company Ltd(西北工业大学; 吉利汽车研究院(宁波)有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出MuSeC音乐语义编解码器,通过分解语义与声学内容,在保证高保真重建的同时生成更利于语言建模的离散标记,提升LLM音乐生成质量。
AI 中文摘要
离散音频分词化已成为近期音乐生成中原始波形与自回归建模之间的关键接口。因此,音乐分词器必须同时支持高保真重建,并产生适合语言建模的离散序列。现有的面向重建的分词器常常将音乐结构与精细声学细节混合,产生高熵且难以建模的标记。相比之下,语义引导的替代方案是为语音设计的,不适合音乐,常常损害重建质量。我们通过围绕可衡量的音乐语义内容概念(基于下游音乐信息检索任务)重新思考音乐分词化来解决这些权衡。在此定义的指导下,我们提出了MuSeC,一种音乐语义编解码器,它直接从混合信号中分解语义和声学内容,无需源分离。MuSeC保留了高保真重建所需的信息,同时产生更有利于语言模型的离散单元。实验上,它提高了重建质量,并产生了更可预测的标记序列,为高保真LLM音乐生成提供了实用基础。演示可在以下网址获取:此https URL。
英文摘要
Discrete audio tokenization has become the critical interface between raw waveforms and autoregressive modeling in recent music generation. As a result, music tokenizers must simultaneously support high-fidelity reconstruction and produce discrete sequences that remain amenable to language modeling. Existing reconstruction-oriented tokenizers often mix musical structure with fine acoustic details, producing high-entropy tokens that are hard to model. In contrast, semantics-guided alternatives are designed for speech and do not fit music well, often hurting reconstruction quality. We address these trade-offs by rethinking music tokenization around a measurable notion of music semantic content grounded in downstream Music Information Retrieval tasks. Guided by this definition, we propose MuSeC, a music semantic codec that factorizes semantic and acoustic content directly from mixed signals without source separation. MuSeC preserves information required for high-fidelity reconstruction while producing more LM-friendly discrete units. Empirically, it improves reconstruction quality and yields more predictable token sequences, providing a practical foundation toward high-fidelity LLM music generation. Demos are available at https://longwaytog0.github.io/MuSeC/.