感知耳语的大语言模型:用于鲁棒耳语语音识别的自监督不确定性学习
Whisper-Aware LLM: Self-Supervised Uncertainty Learning for Robust Whispered Speech Recognition
浏览论文内容
中文总结 AI 辅助
本文提出Whisper-Aware LLM框架,通过自监督学习量化声学信号缺陷并结合置信融合解码机制,在AISHELL6-Whisper数据集上将耳语语音识别的CER相对降低17%,幻觉率降至4.5%,达到最优性能。
中文摘要 AI 辅助
耳语语音的信号歧义性使自动语音识别(ASR)系统出现两种对立的失效模式:要么无法捕捉耳语语音,要么对噪声产生幻觉式转录。本文提出Whisper-Aware LLM(感知耳语的大语言模型)框架,该框架指导音频大语言模型感知并应对这种不确定性。模型通过针对性自监督任务学习量化声学信号的物理缺陷,从而形成内在自我意识;随后通过新颖的Confidence-Fused Decoding(置信融合解码)机制将所学不确定性付诸应用,该机制为大语言模型解码器提供高层指令和帧级注意力调制。实验证实了该方法的有效性:模型在耳语语音识别上达到新的State-of-the-Art(SOTA,最优水平),在AISHELL6-Whisper数据集上字符错误率(CER)相对降低17%;同时直接解决了可靠性权衡问题,幻觉率从25%以上降至4.5%。
英文摘要
The signal ambiguity of whispered speech drives ASR systems toward two opposing failure modes: failing to capture whispered speech or hallucinatory transcription of noise. This paper introduces the Whisper-Aware LLM, a framework that teaches an Audio-LLM to perceive and react to this uncertainty. Our model develops an intrinsic self-awareness by learning to quantify the physical deficiencies of acoustic signals through targeted self-supervised tasks. This learned uncertainty is then operationalized via a novel Confidence-Fused Decoding mechanism, which provides both high-level instructions and frame-level attention modulation to the LLM decoder. Our experiments confirm the effectiveness of this approach. The model sets a new state-of-the-art on whispered speech with a 17% relative CER reduction on AISHELL6-Whisper. At the same time, it directly addresses the reliability trade-off, with hallucination rates dropping from over 25% to 4.5%.
发表机构
- Qwen Business Unit of Alibaba(阿里巴巴通义千问业务部)
机构由 AI 辅助整理,请以论文原文为准。