理解ECAPA-TDNN嵌入的超球面几何及其对零样本语音转换的影响
Understanding Hyperspherical Geometry of ECAPA-TDNN Embedding and Its Impact on Zero-Shot Voice Conversion
浏览论文内容
中文总结 AI 辅助
本文分析ECAPA-TDNN嵌入的超球面几何,提出两种几何正则化策略改善原型分布,提升零样本语音转换对未见说话人的鲁棒性。
中文摘要 AI 辅助
角度间隔说话人编码器广泛用于语音转换,但其分类器原型的几何结构仍鲜为人知。我们将ECAPA-TDNN分类器原型视为单位超球面上的点,并使用旋转不变的角统计量以及全局和局部有效维度度量来表征其组织方式。我们的分析表明,标准训练可能引发角度集中和有效维度的显著降低。为解决这一问题,我们研究了两种几何正则化策略(铰链Riesz对数能量和有效维度最大化),应用于分类器原型以鼓励更均匀的超球面覆盖。所得原型集表现出更高的有效维度和改善的各向同性,且对说话人识别性能的影响取决于配置。当相应的ECAPA-TDNN模型用作Fast-VGAN的说话人编码器时,正则化系统在零样本语音转换中也表现出更好的鲁棒性,尤其是对于未见过的说话人。
英文摘要
Angular-margin speaker encoders are widely used in voice conversion, yet the geometry of their classifier prototypes remains poorly understood. We analyze ECAPA-TDNN classifier prototypes as points on the unit hypersphere and characterize their organization using rotation-invariant angular statistics together with global and local effective dimensionality measures. Our analysis shows that standard training can induce angular concentration and a substantial reduction in effective dimensionality. To address this, we investigate two geometric regularization strategies (hinged Riesz log-energy and effective-dimension maximization) applied to classifier prototypes to encourage more uniform hyperspherical coverage. The resulting prototype sets exhibit higher effective dimensionality and improved isotropy, with configuration-dependent effects on speaker-recognition performance. When the corresponding ECAPA-TDNN models are used as speaker encoders for Fast-VGAN, the regularized systems also exhibit improved robustness in zero-shot voice conversion, particularly for previously unseen speakers.
发表机构
- STMS Lab, IRCAM CNRS, Sorbonne Université(STMS实验室,法国国家科学研究中心下属的IRCAM,索邦大学)
机构由 AI 辅助整理,请以论文原文为准。