arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

语音增强在何处损害识别:推理时间极坐标投影诊断

Where Speech Enhancement Hurts Recognition: An Inference Time Polar Projection Diagnosis

Mingyue Huo, Yuheng Zhang, Hao Zhang

arXiv 2607.11157首次发表:更新:

发表机构

University of Illinois Urbana-Champaign; Wuhan University(伊利诺伊大学厄巴纳-香槟分校; 武汉大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究语音增强对自动语音识别的影响,提出推理时间极坐标投影诊断方法,通过扫描控制参数将ASR退化转化为幅度和相位效应,分析得出幅度强度是关键,最佳强度因识别器而异,该投影可为SE前端提供简单缓解方法。

AI 中文摘要

语音增强(SE)可显著提升感知质量,但增强后的语音不一定能提高自动语音识别(ASR)。现有补救措施可缓解这种不匹配,但常见解释仍停留在定性层面。我们提出推理时间极坐标投影,用于STFT域增强诊断。给定掩码\(M = Ae^{j\phi}\),极坐标投影形成\(M_{\alpha,\gamma}=A^\alpha e^{j\gamma\phi}\)。通过在冻结的SE和ASR模型上扫描这些控制参数,将ASR退化转化为可测量的幅度和相位效应。投影分析表明幅度强度是关键因素,估计的相位校正并无识别优势。最佳幅度强度取决于识别器,波形输入的wav2vec2.0偏好强校正,而对数梅尔输入、抗噪的Whisper则偏好较弱校正。最后,该投影为STFT掩码域中的任何SE前端提供了一种无需重新训练增强器或识别器的简单缓解方法,对依赖增强语音的语音助手和智能体直接有用。

英文摘要

Speech enhancement (SE) can substantially improve perceptual quality, yet enhanced speech does not necessarily improve automatic speech recognition (ASR). Common explanations such as enhancement artifacts and over-suppression remain qualitative and do not localize which enhancement component affects recognition. We study inference-time polar projection, which transforms an STFT mask $M=Ae^{jϕ}$ into $M_{α,γ}=A^αe^{jγϕ}$ and independently varies magnitude strength and estimated phase correction with frozen SE and ASR models. The resulting response curves separate effects that are coupled in ordinary enhanced-vs-noisy comparisons. On VoiceBank+DEMAND with FRCRN, wav2vec2.0 improves from 17.13 to 10.20 WER near full magnitude correction ($α=.85$), while Whisper reaches its best WER under much weaker correction, improving from 6.44 to 5.57 at $α=.25$ and returning to 6.46 at full magnitude. A second enhancer, a robustness-oriented ASR model, and a reverberant WSJ0 mixture set generalize the pattern and show that the preferred magnitude region depends on the recognizer and acoustic condition. Estimated phase correction is not a stable ASR lever in these evaluations: its effect is smaller, setting-dependent, and provides no consistent gain once magnitude strength is fixed.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑