arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.13046eess.AS

基于语音基础模型表示的距离度量的客观可懂度预测

Objective Intelligibility Prediction Using Distance Metrics on Speech Foundation Model Representations

Lyonel Behringer, Andreas Brendel

AI总结:

本文评估预训练语音基础模型嵌入距离在无需训练下预测语音可懂度的效果,发现Whisper模型最后层结合Fréchet距离最佳,优于经典指标且随模型增大而提升。

AI中文摘要:

预训练语音基础模型的高维表示已被证明对客观语音质量和可懂度预测有益。虽然现有的神经可懂度预测工作通常利用此类表示进行任务特定的微调,但在这项工作中,我们评估了此类表示在无需进一步训练的情况下用于可懂度预测的有效性。我们对多个语音基础模型进行了逐层分析,将各种嵌入距离与主观可懂度评分进行相关性分析。结果表明,从Whisper语音识别模型提取的嵌入最为适用,当使用Fréchet音频距离时,最后的编码器和解码器层产生了最佳相关性。值得注意的是,所评估的距离优于经典的可懂度度量,并且比词错误率和字符错误率更为稳健。此外,随着用于提取嵌入的Whisper模型规模的增大,相关性也得到改善。

英文摘要:

High-dimensional representations of pretrained speech foundation models have proven beneficial for objective speech quality and intelligibility prediction. While existing work on neural intelligibility prediction usually leverages such representations for task-specific fine-tuning, in this work we evaluate the usefulness of such representations for intelligibility prediction without any further training. We conduct a layer-wise analysis of multiple speech foundation models, correlating various embedding distances with subjective intelligibility scores. The results show that embeddings extracted from Whisper speech recognition models are best suited, with the last encoder and decoder layers yielding the best correlations when using the Fréchet Audio Distance. Notably, the evaluated distances outperform classical intelligibility metrics and are more robust than Word and Character Error Rates. Further, correlations improve with increasing size of the Whisper model from which embeddings are extracted.

补充信息

↑