arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VOSSA:面向流式语音架构的声纹优化

VOSSA: Voiceprint Optimization for Streaming Speech Architectures

Mu-Ruei Tseng, Waris Quamer, Ghady Nasrallah, Ricardo Gutierrez-Osuna

arXiv 2609.38887首次发表:更新:

发表机构

Department of Computer Science \& Engineering, Texas A\&M University, College Station, US

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对流式语音转换中预训练说话人嵌入与帧级声学生成的冲突,提出VOSSA框架,从内容编码器中间层提取说话人信息并联合训练,在六个数据集上改善F0动态与元音线索,同时保持质量指标并提升感知效果。

AI 中文摘要

实时语音转换(VC)系统通常依赖从自动说话人验证(ASV)模型中预训练的说话人嵌入。虽然这些嵌入在说话人区分方面有效,但它们被训练为在说话人内部的语音和韵律变化中保持稳定,这可能与流式约束下的帧级声学生成相冲突。为解决这一问题,我们提出VOSSA(面向流式语音架构的声纹优化),一种说话人表示框架,从中间内容编码器层提取说话人信息,并使用注意力统计池化进行聚合。该嵌入与VC目标联合训练,无需单独的说话人编码器。在六个数据集上,VOSSA改善了F0动态和元音判别性声学线索,同时保持了相当的NISQA-MOS、WER和说话人相似度。感知测试进一步表明在自然度、说话人相似度、可懂度和活力方面有所提升。

英文摘要

Real-time voice conversion (VC) systems commonly rely on pretrained speaker embeddings from automatic speaker verification (ASV) models. While effective for speaker discrimination, these embeddings are trained to remain stable across phonetic and prosodic variations within-speaker, which may conflict with frame-level acoustic generation in streaming constraints. To address this issue, we propose VOSSA (Voiceprint Optimization for Streaming Speech Architectures), a speaker representation framework that extracts speaker information from intermediate content encoder layers and aggregates using attentive statistics pooling. The embedding is trained jointly with VC objectives, removing the need for a separate speaker encoder. Across six datasets, VOSSA improves F0 dynamics and vowel-discriminative acoustic cues while maintaining comparable NISQA-MOS, WER, and speaker similarity. Perceptual tests further indicate improvements in naturalness, speaker similarity, intelligibility, and vibrancy.

CommentsPublished in Proceedings of Interspeech 2026. Best Student Paper Award

Journal refTseng, M.-R., Quamer, W., Nasrallah, G., Gutierrez-Osuna, R. (2026) VOSSA: Voiceprint Optimization for Streaming Speech Architectures. Proc. Interspeech 2026, 4716-4720

DOI:10.21437/Interspeech.2026-2763

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑