arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33999eess.AScs.SD

重新思考自动语音相似性:从EER转向嵌入几何

Rethinking Automated Voice Similarity by Shifting from EER to Embedding Geometry

Szu-Chi Chen, Jia-Kai Dong, Yi-Cheng Lin, Sung-Feng Huang, Hung-yi Lee

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过人类感知对齐度量发现,说话人验证模型的感知对齐度主要由学习目标而非EER决定,并揭示嵌入几何中有效维度与感知对齐的强负相关,通过维度瓶颈可显著提升对齐度,为语音相似性评估提供几何标准。

中文摘要 AI 辅助

说话人验证(SV)模型通常被假定为随着验证准确率的提高能更好地捕捉说话人特征之间的细微差别,这导致它们在语音生成任务中被广泛用作人类语音相似性的自动化代理。然而,通过建立人类感知对齐度量并进行系统分析,我们证明感知对齐更多地由模型的训练方式(其学习目标)决定,而非其性能表现(EER)。值得注意的是,标准的基于边界的分类损失(如AAM-Softmax)产生的感知对齐度显著低于原型度量损失,而EER本身无法跟踪人类判断,这直接挑战了社区的隐含假设。我们将这种差异追溯到嵌入几何,其中模型的有效维度($d_{\mathrm{eff}}$)与感知对齐度的秩相关系数为$-0.95$,揭示出分类损失偏好的维度扩展与人类语音感知的低维本质根本冲突。施加维度瓶颈可将$d_{\mathrm{eff}}$压缩,并将感知对齐度($\rho_{\mathrm{align}}$)从0.08提升至0.74,从而为评估语音相似性建立了原则性的几何标准。

英文摘要

Speaker verification (SV) models are commonly assumed to better capture nuances among speaker characteristics as verification accuracy improves, leading to their widespread use as automated proxies for human voice similarity in speech generation tasks. However, by establishing a human perceptual alignment metric and conducting systematic analysis, we demonstrate that perceptual alignment is governed far more by how a model is trained (its learning objective) than by how well it performs (EER). Notably, standard margin-based classification losses (e.g., AAM-Softmax) yield substantially lower perceptual alignment than prototypical metric losses, while EER itself fails to track human judgment, directly challenging the community's implicit assumption. We trace this divergence to embedding geometry, where a model's effective dimensionality ($d_{\mathrm{eff}}$) tracks perceptual alignment with a $-0.95$ rank correlation, revealing that the dimensional spread favored by classification losses fundamentally clashes with the low-dimensional nature of human voice perception. Imposing a dimensionality bottleneck compresses $d_{\mathrm{eff}}$ and raises perceptual alignment ($ρ_{\mathrm{align}}$) from 0.08 to 0.74, establishing a principled geometric criterion for evaluating voice similarity.

发表机构

  • National Taiwan University(国立台湾大学)
  • NVIDIA(英伟达)
  • Artificial Intelligence Center of Research Excellence (NTU AI-CoRE)(人工智能卓越研究中心(台大AI-CoRE))

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑