AI 中文总结
研究多话语、多模态环境下的说话人验证,通过对多个匿名话语的音频、韵律和语言信息进行聚合,比较不同聚合策略,发现多模态系统性能更优,帧级聚合EER最低,即便少量匿名话语结合音频和文本也能显著降低EER。
AI 中文摘要
大多数自动说话人验证(ASV)系统针对单个话语运行,而现实世界交互通常包含多个话语。随着语音积累,通过声学、韵律和语言线索可获得越来越丰富的说话人信息,这可能挑战主要针对语音特征的说话人匿名化方法。我们在多话语、多模态环境中研究ASV,探讨跨匿名语音聚合信息是否影响隐私。先研究跨多个匿名话语的纯音频聚合,发现随着语音增多性能持续提升;接着纳入韵律和语言信息,表明多模态系统优于单模态方法;最后比较聚合策略,发现帧级聚合EER最低。即使仅五个匿名话语,结合音频和文本相对于纯音频聚合EER降低超15%,表明匿名化后仍有大量说话人鉴别信息可获取。
英文摘要
Most automatic speaker verification (ASV) systems operate on individual utterances, despite real-world interactions typically consisting of multiple utterances. As speech accumulates, increasingly rich speaker information becomes available through acoustic, prosodic, and linguistic cues, potentially challenging speaker anonymization methods that primarily target vocal characteristics. We investigate ASV in a multi-utterance, multimodal setting and examine whether aggregating information across anonymized speech impacts privacy. We first study audio-only aggregation across multiple anonymized utterances and observe consistent performance improvements as more speech becomes available. We then incorporate prosodic and linguistic information, showing that multimodal systems outperform unimodal approaches. Finally, we compare aggregation strategies and find that frame-level aggregation yields the lowest EERs. Even with only five anonymized utterances, combining audio and text reduces EER by over 15% relative to audio-only aggregation, demonstrating that substantial speaker-discriminative information remains accessible despite anonymization.