PTC-Bias:语音大语言模型中基于音素级时间竞争的偏置检索与解码后校正
PTC-Bias: Phoneme-Level Temporal Competition for Bias Retrieval and Post-Decoding Correction in Speech LLMs
浏览论文内容
中文总结 AI 辅助
提出PTC-Bias,通过音素级时间竞争的两阶段框架(检索与解码后校正),有效利用大规模偏置列表,在LibriSpeech上显著降低B-WER且不损害U-WER。
中文摘要 AI 辅助
上下文偏置能提升语音大语言模型(SpeechLLMs)中的罕见词识别,但高效利用大规模偏置列表仍具挑战。我们提出PTC-Bias,一种基于音素级时间竞争的两阶段框架。在预填充阶段,PTC检索执行帧同步音素解码及候选发音间的时间竞争,生成紧凑的偏置词短列表及对应语音区间。在SpeechLLM解码后,PTC校正对这些区间内检索到的候选词与不匹配的转录片段进行第二次局部竞争。选择性校正减少了近同音词和分词错误,同时保留正确的转录。两个阶段共享相同的音素后验,无需额外的SpeechLLM前向传播。在LibriSpeech上的实验表明,在两种SpeechLLM和高达2000词的偏置列表上均取得一致改进。使用Prompt-SLAM-ASR-7B和2000个偏置词时,PTC-Bias在test-clean/test-other上将B-WER相对CTC-Filter降低了23.4%/23.9%,同时保持U-WER几乎不变。
英文摘要
Contextual biasing improves rare-word recognition in speech large language models (SpeechLLMs), but efficiently exploiting large bias lists remains challenging. We propose PTC-Bias, a two-stage framework based on phoneme-level temporal competition. At the prefill stage, PTC Retrieval performs frame-synchronous phoneme decoding and temporal competition among candidate pronunciations, producing a compact bias-word shortlist and corresponding speech intervals. After SpeechLLM decoding, PTC Correction conducts a second local competition between the retrieved candidates and mismatched transcript spans within these intervals. Selective correction reduces near-homophone and word-segmentation errors while preserving correct transcriptions. Both stages share the same phoneme posteriors and require no additional SpeechLLM forward pass. Experiments on LibriSpeech show consistent gains across two SpeechLLMs and bias lists of up to 2000 words. With Prompt-SLAM-ASR-7B and 2000 bias words, PTC-Bias reduces B-WER by 23.4%/23.9% relative to CTC-Filter on test-clean/test-other, while keeping U-WER nearly unchanged.
发表机构
- Shanghai University(上海大学)
机构由 AI 辅助整理,请以论文原文为准。