arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.20706cs.SDcs.HC

使用混合语音特征提高ASV系统的性能

Improving the performance of an ASV system using hybrid speech features

Stanisław Ciszkiewicz, Artur Janicki

首次发表
浏览论文内容

中文总结 AI 辅助

研究旨在通过结合不同信号表示的混合特征集提升ASV系统性能,从MFCC、CQCC到RAB描述符,在谷歌语音命令数据集上于干净及有噪声场景实验,结果显示混合特征集(PNCC+RAB)能提高有噪声时的说话人验证性能。

中文摘要 AI 辅助

对安全便捷认证方法需求日增,生物识别解决方案愈发流行,基于语音的自动说话人验证(ASV)系统也被采用。但该系统易受多种攻击和声学噪声影响,降低验证准确性。本文研究通过使用混合特征集(从常用的梅尔频率倒谱系数(MFCC)、恒定Q倒谱系数(CQCC)到创新的RAB描述符)提高ASV系统性能的潜力。在谷歌语音命令数据集的录音上,于干净和有噪声两种场景下实验,用EER指标比较系统性能。结果表明,使用混合特征集(PNCC+RAB)可提高有噪声条件下的说话人验证性能。

英文摘要

The growing need for secure and convenient authentication methods has led to the increasing popularity of biometric solutions. In addition to traditional and popular methods, such as fingerprint or iris scanning, voice-based approaches are also employed. User identity verification based on voice is conducted using Automatic Speaker Verification (ASV) systems. Despite their many advantages, these systems are sensitive to various types of attacks and acoustic noises, which can reduce verification accuracy. This work examines the potential to improve the performance of ASV systems by using hybrid feature sets that combine different signal representations, starting with widely-used Mel-Frequency Cepstral Coefficients (MFCC), through Constant Q Cepstral Coefficients (CQCC) and ending with the innovative RAB descriptor. Experiments were conducted on recordings from the Google Speech Commands dataset under two scenarios: in clean conditions and in the presence of acoustic noise. Finally, the systems' performance was compared using the EER metric to determine whether hybrid feature sets decrease verification error. The results show that using a hybrid feature set (PNCC+RAB) improves speaker verification performance under noisy conditions.

↑