arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31652cs.SDcs.AIeess.AS

Open-Qwen-Music:基于LLM的音乐创作与扩散渲染的可审计框架

Open-Qwen-Music: An Auditable Framework for LLM-Based Music Composition and Diffusion Rendering

Yangbin Yu, Mingyu Yang

AI总结:

Open-Qwen-Music是一个完全开源的文本到音乐生成系统,结合LLM语义创作与扩散渲染,提供完整训练数据、流程和预训练权重,作为可复现的研究基线。

AI中文摘要:

我们提出了Open-Qwen-Music,这是对Qwen-Music的开源重构,也是一个完全规范化的文本到音乐生成研究系统,它将基于LLM的语义创作与基于扩散的声学渲染相结合。该系统包含一个25 Hz单码本音乐分词器、一个3B参数的自回归音乐LLM,以及一个生成48 kHz立体声音频的扩散渲染器,遵循Qwen-Music报告的跨模块接口。该设计的最强系统仍然保持封闭,而著名的开源音乐生成项目仅发布权重和推理代码,不包含其训练语料库或端到端训练实现。这限制了对信息损失和预测错误如何从语义表示通过自回归规划传播到声学渲染的独立和受控研究。据我们所知,Open-Qwen-Music是首个完全开源的LLM创作加扩散渲染的文本到音乐生成系统。除了模型权重和推理代码外,发布内容还包括训练数据集和来源清单、完整的数据处理、标注、训练、推理和评估流程、配置以及每个学习模块的预训练权重。工件清单将整个工作流程中这些工件的身份绑定在一起。这些工件共同建立了模块化架构的可复现实现,并为组件级分析和未来评估提供了实证基础。我们将该系统作为一个透明、可执行的研究基线和社区的起点,而不是作为与Qwen-Music质量对等的证据。Open-Qwen-Music是一项持续进行的工作,我们将继续改进其生成质量、可控性和鲁棒性。所有发布工件均可在此https URL获取。

英文摘要:

We present Open-Qwen-Music, an open reconstruction of Qwen-Music and a fully specified research system for text-to-music generation that couples LLM-based semantic composition with diffusion-based acoustic rendering. The system comprises a 25 Hz single-codebook music tokenizer, a 3B-parameter autoregressive Music LLM, and a diffusion renderer producing 48 kHz stereo audio, following the cross-module interfaces reported by Qwen-Music. The strongest systems of this design remain closed, and prominent open music-generation projects release weights and inference code without their training corpora or end-to-end training implementations. This limits independent and controlled study of how information loss and prediction errors propagate from semantic representation through autoregressive planning to acoustic rendering. To our knowledge, Open-Qwen-Music is the first fully open release of an LLM-composition-plus-diffusion-rendering text-to-music system. Beyond model weights and inference code, the release includes the training datasets and provenance manifests, complete data-processing, annotation, training, inference, and evaluation pipelines, configurations, and pretrained weights for every learned module. Artifact manifests bind the identities of these artifacts across the complete workflow. Together, these artifacts establish a reproducible implementation of the modular architecture and provide an empirical basis for component-level analysis and future evaluation. We present the system as a transparent, executable research baseline and a starting point for the community, not as evidence of quality parity with Qwen-Music. Open-Qwen-Music is an ongoing effort, and we will continue to improve its generation quality, controllability, and robustness. All release artifacts are available at https://github.com/biang15343100-source/Open-Qwen-Music.

↑