CORD-KWS:面向开放词汇关键词识别的校准与顺序感知检测
CORD-KWS: Calibrated, Order-Aware Detection for Open-Vocabulary Keyword Spotting
浏览论文内容
中文总结 AI 辅助
CORD-KWS提出一种训练框架,通过校准检测头和帧级CTC目标,在保持嵌入模型效率的同时提升开放词汇关键词识别性能,在LibriPhrase基准上超越现有方法。
中文摘要 AI 辅助
开放词汇关键词识别(KWS)必须在不重新训练的情况下检测任意关键词。基于交叉注意力的模型达到了最先进的性能,但需要音频与每个关键词之间的成对交互,为每个音频-关键词对重新计算融合表示。基于嵌入的模型通过独立编码和相似度评分避免了这种计算成本,但在具有挑战性的基准(如LibriPhrase-hard)上仍然表现不佳。我们表明,这一差距主要是训练问题,而非模型容量的限制。虽然对比学习鼓励正确关键词排在负样本之上,但它并未显式校准绝对相似度分数,而绝对相似度分数对于在关键词检测中应用固定阈值至关重要。我们提出了校准的顺序感知检测KWS(CORD-KWS),这是一个训练框架,在保留相同编码器、余弦评分和O(1)关键词注册的同时,引入了两个互补目标:一个校准检测头用于监督绝对分数,以及一个帧级CTC目标用于保留音素顺序信息。CORD-KWS在LibriPhrase-easy和LibriPhrase-hard上分别实现了0.43%和9.64%的等错误率(EER),优于现有的基于嵌入和基于交叉注意力的模型。
英文摘要
Open-vocabulary keyword spotting (KWS) must detect arbitrary keywords without retraining. Cross-attention-based models achieve state-of-the-art performance but require pairwise interaction between the audio and each keyword, recomputing the fused representation for every audio--keyword pair. Embedding-based models avoid this computational cost through independent encoding and similarity scoring, yet remain inferior on challenging benchmarks such as LibriPhrase-hard. We show that this gap is primarily a training issue rather than a limitation of model capacity. While contrastive learning encourages correct keywords to rank above negatives, it does not explicitly calibrate absolute similarity scores, which are critical for applying a fixed threshold across keyword detection. We propose Calibrated, Order-aware Detection KWS (CORD-KWS), a training framework that introduces two complementary objectives while retaining the same encoders, cosine scoring, and O(1) keyword enrollment: a calibrated detection head that supervises absolute scores, and a frame-level CTC objective that preserves phonetic order information. CORD-KWS achieves EERs of 0.43% and 9.64% on LibriPhrase-easy and LibriPhrase-hard, respectively, outperforming existing embedding- and cross-attention-based models.