发表机构
Ant Group(蚂蚁集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出MECT,一种集成专家混合机制到CNN-Transformer骨干的说话人验证模型,通过优化块结构和路由策略,在保持紧凑参数的同时,在VoxCeleb1和CN-Celeb上取得领先性能,并支持流式推理。
AI 中文摘要
本文提出了MECT,一种说话人验证模型,该模型将专家混合(MoE)机制集成到具有优化块结构和堆叠方案的CNN-Transformer骨干网络中。具体而言,我们研究了四种MoE变体,它们跨越话语级和帧级粒度,并采用密集和稀疏路由策略。与没有MoE的基线相比,MoE机制被证明是有效的,且仅增加了少量参数。我们进一步将MECT扩展到一系列模型规模,所有模型均保持紧凑的参数和较低的计算复杂度。特别是,MECT-B2在VoxCeleb1上达到了最先进的性能,并在CN-Celeb上取得了强劲的结果,展示了其在不同数据集上的有效性。此外,我们通过因果重训练建立了一种流式推理范式,该范式在100毫秒的块大小下保持了强劲的性能。
英文摘要
In this paper, we propose MECT, a speaker verification model that integrates the Mixture-of-Experts (MoE) mechanism into a CNN-Transformer backbone with optimized block structure and stacking scheme. Specifically, we investigated four MoE variants that span utterance-level and frame-level granularity with dense and sparse routing strategies. The MoE mechanism proves to be effective over the baseline without MoE with only a small increase in parameters. We further scale MECT to a series of model sizes, all maintaining compact parameters and low computational complexity. In particular, MECT-B2 achieves state-of-the-art performance on VoxCeleb1 and delivers strong results on CN-Celeb, demonstrating its effectiveness across diverse datasets. In addition, we establish a streaming inference paradigm through causal retraining, which maintains strong performance at a chunk size of 100ms.
Comments5 pages