AI 中文总结
针对现有音频系统难以组织多类型音频形成长时场景的问题,提出 Qwen-Audio-3.0-Gen-Preview 框架,采用 DiT 与共享 VAE 实现统一多域音频生成,在多个基准上展现出性能优势。
AI 中文摘要
现有的单域及多任务音频系统在直接将语音、音乐、音效、环境音及多个角色组织成长时长时间场景方面仍存在局限。我们提出 Qwen-Audio-3.0-Gen-Preview,这是一个统一的非自回归框架,采用扩散 Transformer(DiT)与共享变分自动编码器(VAE)生成完整混合波形。提示增强将自由形式请求转换为结构化时间记录,渲染为文本条件;两阶段数据课程与语义条件视图训练模型,使其可跨域使用这些条件。共享连续 VAE 将 48kHz 立体声波形压缩为 25Hz 潜序列,并融入语义监督,为语音、音乐、音效及其混合提供单一表示。在 Seed-TTS-Eval 上,该模型在所有三个子集的说话人相似度表现为最突出优势;在多说话人基准上,该模型在两种语言中均比 Seed-Audio-1.0 展现出更高的跨轮一致性。在 AudioCaps 上,其优势集中于使用大型音频语言模型与 AudioBox 的评估中;相比 Seed-Audio-1.0,它实现了更强的时间定位能力。使用约 10% 专有内部模型的音乐数据,该模型在 SongBench 的全部七个组件中表现接近,其中三个组件领先,同时保留了语音与通用音频能力。这些结果证明了统一生成在时间结构化多域音频领域的潜力。
英文摘要
Existing single-domain and multi-task audio systems remain limited in directly organizing heterogeneous audio components, ambience, and multiple roles into long-form temporal scenes. We present Qwen-Audio-3.0-Gen-Preview, a unified non-autoregressive framework that uses a Diffusion Transformer (DiT) and a shared variational autoencoder (VAE) to generate the complete mixed waveform. Prompt enhancement converts free-form requests into structured temporal records that are rendered as textual conditions, while a two-stage data curriculum and semantic conditional views train the proposed model to use these conditions across standalone and mixed-scene audio. A shared continuous VAE compresses 48kHz stereo waveforms into 25Hz latent sequences and incorporates semantic supervision, providing one representation for heterogeneous audio. On the public reference-conditioned benchmark, speaker similarity is the proposed model's clearest strength across all three subsets. Across the multi-speaker and rich-timeline benchmarks, its clearest comparative strengths are cross-turn consistency in both languages and temporal localization, respectively. On AudioCaps, its advantages are concentrated in evaluations using large audio-language models and AudioBox. These results demonstrate the potential of unified generation for temporally structured audio without task-specific branches.