arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

神经音频编解码器对非洲语音的鲁棒性如何?多任务基准与感知质量的局限性

How Robust Are Neural Audio Codecs for African Speech? A Multi-Task Benchmark and the Limits of Perceptual Quality

Chibuzor Okocha, Christan Earl Grant

arXiv 2610.00154首次发表:更新:

发表机构

University of Florida(佛罗里达大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过多任务基准测试七个神经音频编解码器在非洲语音上的鲁棒性,发现信号指标与下游退化相关性差异大,且退化可通过LoRA适配部分恢复,强调任务感知评估的必要性。

AI 中文摘要

神经音频编解码器实现了低比特率语音压缩和分词化,但其在带口音和多语言语音上的鲁棒性仍未得到充分评估。我们在三个非洲语音数据集(afrinames、afrispeech dialog、afrispeech multilingual)上对七个开源神经编解码器(DAC、EnCodec、FocalCodec、LanguageCodec、SemantiCodec、UniCodec、WavTokenizer)进行了基准测试,报告了信号级质量(NISQA、UTMOS、ViSQOL、STOI、F0-RMSE)以及两个下游任务:自动语音识别(ASR)和说话人验证(ASV)。我们的分析得出四个发现。第一,信号指标在下游有效性上差异显著:基于参考的结构/可懂度度量(ViSQOL、STOI)和韵律误差(F0-RMSE)追踪ASR和ASV退化远比神经平均意见得分预测器(NISQA、UTMOS)可靠。第二,可懂度和说话人身份保留在不同架构间差异显著,且表观身份排名本身依赖于ASV后端。第三,退化强烈依赖于领域,且在对话式对话中最为严重。第四,由此产生的退化部分可恢复:在大约35小时的非洲语音上进行的参数高效编解码器适配(LoRA,约1-3%的参数)缩小了压缩引起的词错误率差距,将识别恢复到未压缩基线水平(在配套工作中开发)。这些结果促使对语音编解码器进行任务感知、领域代表性和适配感知的评估,作为包容性部署的前提。

英文摘要

Neural audio codecs enable low-bitrate speech compression and tokenization, yet their robustness on accented and multilingual speech remains under-evaluated. We benchmark seven open-source neural codecs (DAC, EnCodec, FocalCodec, LanguageCodec, SemantiCodec, UniCodec, WavTokenizer) on three African speech datasets afrinames, afrispeech dialog, afrispeech multilingual, reporting signal-level quality (NISQA, UTMOS, ViSQOL, STOI, F0-RMSE) alongside two downstream tasks: automatic speech recognition (ASR) and speaker verification (ASV). Our analysis yields four findings. First, signal metrics differ sharply in downstream validity: reference-based structural/intelligibility measures (ViSQOL, STOI) and prosodic error (F0-RMSE) track ASR and ASV degradation far more reliably than neural mean-opinion-score predictors (NISQA, UTMOS). Second, intelligibility and speaker-identity preservation diverge sharply across architectures, and the apparent identity ranking itself depends on the ASV backend. Third, degradation is strongly domain-dependent and largest for conversational dialog. Fourth, the resulting degradation is partly recoverable: parameter-efficient \emph{codec} adaptation (LoRA, $\sim$1--3\% of parameters) on roughly 35 hours of African speech reduces the compression-induced word-error-rate gap, recovering recognition toward the uncompressed baseline (developed in companion work). These results motivate task-aware, domain-representative, and adaptation-aware evaluation of speech codecs as a prerequisite for inclusive deployment.

CommentsAccepted to IEEE Speech Language Technology

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑