arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.11804eess.AScs.SD

MiDashengLM-Gen:基于大语言模型驱动的自回归流匹配的统一音频场景生成

MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching

Xingwei Sun, Heinrich Dinkel, Gang Li, Jiahao Mei, Yadong Niu, Zerui Han, Yuepeng Jiang, Jiahao Zhou, Lichun Fan, Jian Luan

首次发表
浏览论文内容

中文总结 AI 辅助

MiDashengLM-Gen是首个端到端训练的通用文本到音频生成模型,通过结合LLM与每token条件流匹配,在Seed-TTS等基准上大幅提升语音可懂度,且支持多语言场景。

中文摘要 AI 辅助

生成同时融合语音、音乐和音效的连贯音频场景仍是一项重大挑战。现有方法通常依赖不连贯的流水线,其中冻结的、解耦的文本编码器为独立的音频解码器提供输入,这限制了跨模态优化并导致语音可懂度较差。为克服这些限制,我们引入MiDashengLM-Gen,这是一种端到端框架,它将预训练的大语言模型(LLM)与每token条件流匹配相结合,用于自回归、可变长度的混合音频场景生成。MiDashengLM-Gen是首个采用单一端到端训练模型进行通用文本到音频生成的方法。实证评估表明,MiDashengLM-Gen与现有统一模型相比大幅提升了语音可懂度:在Seed-TTS基准上,英文词错误率(WER)从12.15%降至2.79%,接近专用文本到语音(TTS)系统的性能(1.24%);此外,该框架可有效扩展至多语言场景,与现有基线相比产生极具竞争力的多语言WER;最后,该模型在MECAT基准上保持了具有竞争力的混合音频生成质量。代码和检查点可在相关链接获取,演示页面也可通过对应链接访问。

英文摘要

Generating coherent audio scenes that simultaneously blend speech, music, and sound effects remains a significant challenge. Current approaches typically rely on a disjointed pipeline where a frozen, decoupled text encoder feeds a separate audio decoder, limiting cross-modal optimization and leading to poor speech intelligibility. To overcome these limitations, we introduce MiDashengLM-Gen, an end-to-end framework that couples a pre-trained Large Language Model (LLM) with per-token conditional flow matching for autoregressive, variable-length mixed-audio scene generation. MiDashengLM-Gen represents a first approach for general text-to-audio generation with one end-to-end trained model. Empirical evaluations demonstrate that MiDashengLM-Gen drastically improves speech intelligibility over existing unified models. On the Seed-TTS benchmark, English Word Error Rate (WER) drops from 12.15% to 2.79%, approaching the performance of dedicated Text-to-Speech (TTS) systems (1.24%). Furthermore, the framework extends effectively to multilingual settings, yielding highly competitive multilingual WERs compared to existing baselines. Lastly, the model maintains competitive mixed-audio generation quality on the MECAT benchmark. Code and checkpoints are available at https://github.com/xiaomi-research/midashenglm-gen and https://huggingface.co/mispeech/midashenglm-gen, and the demo page is available at https://xingws.github.io/midashenglm-gen-demo/.

发表机构

  • MiLM Plus
  • Xiaomi Inc.(小米公司)
  • Shanghai Jiao Tong University(上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

↑