arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用语义锚定语音:一种针对低资源语言自动语音识别的多模态适配器机制

Anchoring Speech with Semantics: A Multimodal Adapter Mechanism for Automatic Speech Recognition in Low-Resource Languages

Kuan-Tang Huang, Cheng-Yeh Yang, Chien-Chun Wang, Hung-Shin Lee, Hsin-Min Wang, Berlin Chen

arXiv 2608.29239首次发表:更新:

发表机构

National Taiwan Normal University; EZAI; Academia Sinica(台湾师范大学; EZAI; 中央研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对低资源语言ASR的监督证据不足问题,提出轻量级多模态适配器机制SAMA-ASR,结合语义与声学锚,在台湾闽南语和客家话数据集上优于多种基线,且紧凑ST模型可生成有效语义锚。

AI 中文摘要

低资源语言的自动语音识别(ASR)仍然存在困难,因为稀缺的转录本为目标端生成提供的监督证据有限。为解决这一差距,我们提出SAMA-ASR,这是一种轻量级适配器机制,它用辅助翻译的语义锚和语音的声学锚增强解码器;原则上,该机制可应用于类似的编码器-解码器多任务语音模型。通过跨模态适配,SAMA-ASR使解码器状态以翻译衍生的语义嵌入和语音嵌入为条件,在 token 预测前将话语级含义与语音基础证据相结合。在评估时,这些语义锚可由上游语音到文本翻译器自动生成,而非作为神谕翻译提供。对两个各30小时的数据集(涵盖低资源汉藏语系变体台湾闽南语和客家话)的实验表明,SAMA-ASR在声学基线、先前基于提示的基线和仅语义的翻译引导基线之上实现了性能提升,且在实际自动语义锚设置中仍有效;翻译器容量分析显示,有用的语义锚可由紧凑的ST模型生成。

英文摘要

Low-resource ASR remains difficult because scarce transcripts provide limited supervised evidence for target-side generation. To address this gap, we propose SAMA-ASR, a lightweight adapter mechanism that augments the decoder with semantic anchors from auxiliary translations and an acoustic anchor from speech; in principle, the mechanism can be applied to similar encoder--decoder multitask speech models. Through cross-modal adaptation, SAMA-ASR conditions decoder states on translation-derived semantic embeddings and a speech embedding, combining utterance-level meaning with speech-grounded evidence before token prediction. At evaluation time, these semantic anchors can be generated automatically by an upstream speech-to-text translator rather than supplied as oracle translations. Experiments on two 30-hour datasets covering the low-resource Sinitic varieties Taiwanese Hokkien and Hakka show that SAMA-ASR improves over acoustic, prior prompt-based, and semantic-only translation-guided baselines and remains effective in practical automatic semantic-anchor settings; translator-capacity analyses show that useful semantic anchors can be produced by a compact ST model.

CommentsAccepted to EMNLP 2026 (Main Conference)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑