AI 中文总结
AuK 是一个开源基础模型,通过统一接口实现语音生成与编辑,采用多模态大语言模型和混合流 Transformer 架构,经后训练优化与蒸馏,实现高效推理并在多项任务上取得领先性能。
AI 中文摘要
我们推出 AuK,一个开源基础模型,通过自然语言指令和音频上下文的统一接口,将语音生成与编辑功能整合于一体。为支撑这一广泛的能力集,我们构建了约 30.3 亿条指令-音频实例,以及涵盖五大任务族(语音生成、内容编辑、增强与分离、副语言编辑、声学编辑)的 195 万小时有效监督数据。AuK 结合了用于语义条件化的多模态大语言模型、在语音、通用音频和音乐上联合训练的 VAE 以进行声学条件化,以及一个混合修正流 Transformer,该 Transformer 执行双流 MMDiT 块,随后是统一的单流 DiT 块用于生成。训练从仅生成的预热开始,然后进行生成-编辑联合预训练。随后我们应用互补的后训练策略:针对开放式编辑的人类反馈偏好优化,以及针对语音生成的基于奖励的强化学习。为降低推理成本,我们进一步通过一致性初始化和任务路由的解耦 DMD 对模型进行蒸馏。由此产生的 AuK-Flash 在无分类器引导的情况下执行 4 步推理,在匹配条件下相比完整模型实现了 4.5 倍的墙钟加速。实验表明,在零样本和指令控制的语音生成以及通用指令引导编辑方面表现出领先性能,同时在信号级恢复任务上保持竞争力。我们发布了源代码和模型权重,以支持可复现性和进一步研究。
英文摘要
We introduce AuK, an open-source foundational model that unifies speech generation and editing through a common interface of natural-language instructions and audio context. To support this broad capability set, we construct approximately 3.03 billion instruction--audio instances and 1.95 million hours of effective supervision across five task families: speech generation, content editing, enhancement and separation, paralinguistic editing, and acoustic editing. AuK combines a multimodal large language model for semantic conditioning, an VAE jointly trained on speech, general audio, and music for acoustic conditioning, and a hybrid rectified-flow Transformer that performs dual-stream MMDiT blocks followed by unified single-stream DiT blocks for generation. Training begins with generation-only warm-up and proceeds to joint generation--editing pre-training. We then apply complementary post-training strategies: human-feedback preference optimization for open-ended editing and reward-based reinforcement learning for speech generation. To reduce inference cost, we further distill the model with consistency initialization and task-routed Decoupled DMD. The resulting AuK-Flash performs 4-step inference without classifier-free guidance and achieves a 4.5 wall-clock speedup over the full model under matched conditions. Experiments demonstrate leading performance on zero-shot and instruction-controlled speech generation and general instruction-guided editing, while remaining competitive on signal-level restoration tasks. We release both the source code and model weights to support reproducibility and further research.
CommentsOpen-source at https://github.com/Tencent-Hunyuan/AuK