发表机构
Oakland University(奥克兰大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
HybridSB-MoE是一种双域语音增强框架,通过非对称不确定性融合、异构MoE路由及离散化界保证,在VoiceBank+DEMAND数据集上实现了优于扩散与SB基线的性能。
AI 中文摘要
生成式语音增强存在三类差距:频谱模型能捕获谐波结构但常破坏相位,波形模型可保留相位但会遗漏谐波,薛定谔桥(SB)能缩短从噪声到纯净语音的传输过程,但推理成本仅与训练松散关联。本文提出HybridSB-MoE,这是一种双域框架,通过统一的非对称设计原则填补上述差距,包含三项贡献:(i)非对称不确定性融合:频谱路径通过专家分歧捕获认知不确定性,波形桥通过随机动力学建模随机方差,二者非对称融合,使混合权重可适应不同误差 regime,而非平均预测;(ii)异构混合专家(MoE),采用跨五种不同架构原型的top-k=2路由,架构多样性让认知信号能指示哪种归纳偏置失效,而非相似专家间的微小扰动;(iii)离散化界(定理1):路径一致性与轨迹正则化项共同将K步桥采样误差在2-沃尔什斯坦因距离上以K^-α速率界定,使小K推理成为目标层面的保证,而非经验性主张。在VoiceBank+DEMAND数据集上,HybridSB-MoE在对应步预算下优于扩散模型及SB基线,且与一致性蒸馏少步方法性能相当。
英文摘要
Generative speech enhancement faces three gaps: spectral models capture harmonic structure but often disrupt phase, waveform models preserve phase but miss harmonics, and Schrödinger Bridges (SB) shorten transport from noise to clean speech but leave inference cost only loosely tied to training. We propose HybridSB-MoE, a dual-domain framework that fills these gaps through three contributions unified by a single asymmetric design principle. (i) Asymmetric uncertainty fusion: The spectral path captures epistemic uncertainty via expert disagreement, while the waveform bridge models aleatoric variance through stochastic dynamics. We fuse them asymmetrically, allowing the mixing weight to adapt to distinct error regimes rather than average predictions. (ii) Heterogeneous MoE with top-k=2 routing across five distinct architectural archetypes, where architectural diversity makes the epistemic signal indicate which inductive bias fails rather than small perturbations among similar experts. (iii) Discretization bound (Theorem 1): path-consistency and trajectory regularizers together bound the K-step bridge sampling error in 2-Wasserstein distance at rate K-alpha, making small-K inference an objective-level guarantee rather than an empirical claim. On VoiceBank+DEMAND, HybridSB-MoE outperforms diffusion- and SB-based baselines at their step budgets while remaining competitive with consistency-distilled few-step methods.