arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SonicWeave:用于统一音频场景生成的分块路由混合专家模型

SonicWeave: Chunk-Routed Mixture-of-Experts for Unified Audio Scene Generation

Yunrui Cai, Xu Li, Yucheng Zhou, Jinchao Li, Dingdong Wang, Dongchao Yang, Xixin Wu, Chen Zhang, Zhiyong Wu, Pengfei Wan, Helen Meng

arXiv 2608.09571首次发表:更新:

发表机构

Kuaishou Technology; The Chinese University of Hong Kong; Shenzhen International Graduate School, Tsinghua University(快手科技; 香港中文大学; 清华大学深圳国际研究生院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出SonicWeave,一种带CPE-MoE的流匹配模型,可通过单一权重集生成含语音、音乐等的统一音频场景,在多项基准及复杂场景评估中表现优于基线。

AI 中文摘要

文本条件下的通用音频生成正从孤立的语音、音乐和音效合成,向可将其组合为可控、连贯音频场景的单一模型发展。这种统一设置极具挑战性:异构组件对共享主干网络施加了相互冲突的结构要求,而复杂混合场景可能包含局部不同或重叠的内容,需在同一段音频内进行细粒度适配。现有的音频混合专家模型(MoE)主要在域级别进行路由,而令牌级路由忽略了声学信号固有的局部连续性。我们提出SonicWeave,一种用于统一音频场景生成的流匹配模型。其核心是带有冲突门控先验证据路由机制(CPE-MoE)的分块路由混合专家模型。CPE-MoE通过结合编码结构化文本条件和扩散阶段的全局先验,以及来自演化声学状态的局部证据,对连续声学分块进行路由。一个学习得到的冲突门在局部状态不可靠时倾向于先验,而当某一区域偏离全局场景上下文时,允许局部证据影响路由。SonicWeave使用单一权重集支持语音、音乐、音效、演唱及其细粒度混合。在TTS、TTA和TTM基准测试中,SonicWeave始终优于受控的Dense和Base-MoE基线。复杂场景评估进一步证明其组合质量得到提升,而路由分析揭示了跨扩散阶段的内容依赖型专家专业化。这些结果表明,时间连贯的先验证据路由是统一音频生成的有效条件计算策略。项目页面:this https URL。

英文摘要

Text-conditioned general audio generation is moving beyond isolated speech, music, and sound-effect synthesis toward a single model that can compose them into controllable, coherent audio scenes. This unified setting is particularly challenging: heterogeneous components impose conflicting structural requirements on a shared backbone, while a complex mixed scene may contain locally distinct or overlapping content that demands fine-grained adaptation within the same clip. Existing audio mixture-of-experts (MoEs) mainly route at the domain level, while token-wise routing overlooks the local continuity inherent to acoustic signals. We propose SonicWeave, a flow-matching model for unified audio scene generation. At its core is a chunk-routed MoE with a conflict-gated prior-evidence routing mechanism (CPE-MoE). CPE-MoE routes contiguous acoustic chunks by combining a global prior that encodes the structured text condition and diffusion phase with local evidence from the evolving acoustic state. A learned conflict gate favors the prior when local states are unreliable, while allowing local evidence to influence routing when a region departs from the global scene context. SonicWeave supports speech, music, sound effects, singing, and their fine-grained mixtures with a single set of weights. Across TTS, TTA, and TTM benchmarks, SonicWeave consistently improves over controlled Dense and Base-MoE baselines. Complex-scene evaluation further demonstrates improved compositional quality, while routing analyses reveal content-dependent expert specialization across diffusion phases. These results suggest that temporally coherent, prior-evidence routing is an effective conditional-computation strategy for unified audio generation. Project page: https://caiyunrui.github.io/SonicWeave.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑