arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

GIVE-KWS:用于噪声鲁棒查询示例关键词识别的视觉证据门控注入

GIVE-KWS: Gated Injection of Visual Evidence for Noise-Robust Query-by-Example Keyword Spotting

Ming-Hsiang Hu, Kuan-Tang Huang, Hung-Shin Lee, Berlin Chen

arXiv 2610.07046首次发表:更新:

发表机构

National Taiwan Normal University; Graduate Institute of AI Interdisciplinary Applied Technology, National Taiwan Normal University(国立台湾师范大学; 国立台湾师范大学人工智慧跨领域应用技术研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对噪声环境下的查询示例关键词识别,提出GIVE-KWS,通过门控交叉注意力注入视觉证据,在-10 dB下将未见关键词等错误率降低72.9%。

AI 中文摘要

视觉语音有望实现噪声鲁棒的关键词识别,然而视觉流并不一定被使用。在三模态查询示例关键词识别(QbyE-KWS)基准上,我们发现,使用任务训练的视觉编码器的系统在-10 dB信噪比下的等错误率(EER)比文本和音频系统仅高2个百分点,并将这一差距归因于编码器缺乏音素信息。我们提出了GIVE-KWS,其融合阶段GIVE(视觉证据门控注入)通过门控交叉注意力将查询音频条件化于唇部运动。我们表明,视觉鲁棒性依赖于两个相互作用的条件:承载音素的视觉表示,以及注入视觉证据而非重新缩放音频特征的融合。在承载音素的编码器下,注入在-10 dB下比掩蔽产生4.0-9.3 dB的有效信噪比增益,而在音素贫乏的编码器下,该增益几乎消失。相对于基准系统,GIVE-KWS在-10 dB下将未见关键词的EER降低了72.9%,平均降低了62.8%。

英文摘要

Visual speech promises noise-robust keyword spotting, yet a visual stream is not necessarily used. On a tri-modal query-by-example keyword spotting (QbyE-KWS) benchmark, we find that a system with a task-trained visual encoder comes within 2 percentage points of a text-and-audio system in equal error rate (EER) at -10 dB, and link this gap to the encoder's lack of phonemic information. We present GIVE-KWS, whose fusion stage, GIVE (Gated Injection of Visual Evidence), conditions query audio on lip motion through gated cross-attention. We show that visual robustness depends on two interacting conditions: a phoneme-bearing visual representation, and fusion that injects visual evidence rather than rescaling audio features. Under a phoneme-bearing encoder, injection yields an effective SNR gain of 4.0-9.3 dB over masking at -10 dB, whereas under a phoneme-poor one it nearly vanishes. Relative to the benchmark system, GIVE-KWS reduces unseen-keyword EER by 72.9% at -10 dB and 62.8% on average.

CommentsSubmitted to ICASSP 2027

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑