arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Qwen-Audio-3.0-Gen-Preview 技术报告

Qwen-Audio-3.0-Gen-Preview Technical Report

Junyu Dai, Xiaoyue Duan, Xinyue Fan, Yihan Feng, Jingbei Li, Xiangang Li, Yunjia Li, Lejun Min, Yufei Shi, Xingchen Song, Yiran Wang, Cheng Wen, Menglin Wu, Bajian Xiang, Huaicheng Zhang, Han Zhao, Ruichen Zheng

arXiv 2607.27011首次发表:更新:

AI 中文总结

针对现有音频系统难以组织多类型音频形成长时场景的问题,提出 Qwen-Audio-3.0-Gen-Preview 框架,采用 DiT 与共享 VAE 实现统一多域音频生成,在多个基准上展现出性能优势。

AI 中文摘要

现有的单域及多任务音频系统在直接将语音、音乐、音效、环境音及多个角色组织成长时长时间场景方面仍存在局限。我们提出 Qwen-Audio-3.0-Gen-Preview,这是一个统一的非自回归框架,采用扩散 Transformer(DiT)与共享变分自动编码器(VAE)生成完整混合波形。提示增强将自由形式请求转换为结构化时间记录,渲染为文本条件;两阶段数据课程与语义条件视图训练模型,使其可跨域使用这些条件。共享连续 VAE 将 48kHz 立体声波形压缩为 25Hz 潜序列,并融入语义监督,为语音、音乐、音效及其混合提供单一表示。在 Seed-TTS-Eval 上,该模型在所有三个子集的说话人相似度表现为最突出优势;在多说话人基准上,该模型在两种语言中均比 Seed-Audio-1.0 展现出更高的跨轮一致性。在 AudioCaps 上,其优势集中于使用大型音频语言模型与 AudioBox 的评估中;相比 Seed-Audio-1.0,它实现了更强的时间定位能力。使用约 10% 专有内部模型的音乐数据,该模型在 SongBench 的全部七个组件中表现接近,其中三个组件领先,同时保留了语音与通用音频能力。这些结果证明了统一生成在时间结构化多域音频领域的潜力。

英文摘要

Existing single-domain and multi-task audio systems remain limited in directly organizing heterogeneous audio components, ambience, and multiple roles into long-form temporal scenes. We present Qwen-Audio-3.0-Gen-Preview, a unified non-autoregressive framework that uses a Diffusion Transformer (DiT) and a shared variational autoencoder (VAE) to generate the complete mixed waveform. Prompt enhancement converts free-form requests into structured temporal records that are rendered as textual conditions, while a two-stage data curriculum and semantic conditional views train the proposed model to use these conditions across standalone and mixed-scene audio. A shared continuous VAE compresses 48kHz stereo waveforms into 25Hz latent sequences and incorporates semantic supervision, providing one representation for heterogeneous audio. On the public reference-conditioned benchmark, speaker similarity is the proposed model's clearest strength across all three subsets. Across the multi-speaker and rich-timeline benchmarks, its clearest comparative strengths are cross-turn consistency in both languages and temporal localization, respectively. On AudioCaps, its advantages are concentrated in evaluations using large audio-language models and AudioBox. These results demonstrate the potential of unified generation for temporally structured audio without task-specific branches.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑