发表机构
Seoul National University; Research Institute of Mathematics(首尔大学; 数学研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一种基于vMF轮廓似然的简单闭式目标,用于无标签说话人嵌入增强,在多个基准上保持基线性能并提升失配场景效果,表明无需高度结构化公式。
AI 中文摘要
嵌入增强在不修改冻结骨干网络的情况下改善声学失配下的说话人验证。近期工作为此任务建立了实用的无标签设置,但往往采用日益结构化的公式。在此,干净目标在训练期间被直接观测,使得增强成为单位超球面上的匹配问题。我们使用von Mises--Fisher(vMF)似然对干净目标建模,并轮廓化样本级浓度参数,得到一个具有自适应权重的简单闭式目标。在VoxCeleb1、VoxSRC23、CN-Celeb、VOiCES和VC-Mix上,所提方法大体上保持基线性能,并在具有挑战性的失配集合上给出更清晰的增益。在广泛的单视角配方下,它保持稳定,而最近的扩散基线在受控比较中变得不太可靠。这些结果表明,在此设置下,有效的无标签嵌入增强不需要高度结构化的公式。
英文摘要
Embedding enhancement improves speaker verification under acoustic mismatch without modifying a frozen backbone. Recent work has established a practical label-free setting for this task, but often adopts increasingly structured formulations. Here, the clean target is directly observed during training, making enhancement a matching problem on the unit hypersphere. We model the clean target with a von Mises--Fisher (vMF) likelihood and profile out a sample-wise concentration parameter, yielding a simple closed-form objective with adaptive weighting. Across VoxCeleb1, VoxSRC23, CN-Celeb, VOiCES, and VC-Mix, the proposed method largely preserves the baseline and gives clearer gains on challenging mismatch sets. It also remains stable under a broad single-view recipe, where a recent diffusion baseline becomes less reliable in controlled comparisons. These results suggest that effective label-free embedding enhancement in this setting does not require a highly structured formulation.
Comments5 pages. Published in Interspeech 2026
Journal refProc. Interspeech 2026, pp. 393-397
DOI:10.21437/Interspeech.2026-1146