arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过模块化网络扩展实现可扩展的关键词识别

Scalable Keyword Spotting via Modular Network Expansion

Viktor Khaymonenko, Dzmitry Saladukha, Aliaksei Rak, Alexander Rostov

arXiv 2607.19918首次发表:更新:

发表机构

Yandex(Yandex)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究嵌入式设备上KWS模型添加新关键词难的问题,提出参数限制的模块化扩展方法,在固定操作点降低新关键词FRR,优于基线且减少MACs,保留现有关键词相关参数。

AI 中文摘要

嵌入式设备上的关键词识别(KWS)模型在部署后通常需要添加新关键词,但在原始训练数据不可用且对现有触发词进行回归不可接受时,更新很困难。在固定操作点,我们的方法将平均新关键词错误拒绝率(FRR)从6.46降至4.37,优于参数匹配的单独模型基线和参数高效调整基线(适配器、LoRA),同时在相同的添加参数预算(≤10k)下使用更少的乘法累加操作(MACs):16.34M对18.45M/20.52M。我们通过参数限制的模块化扩展实现了这一点:基础网络,包括批归一化统计和核心分类器,被冻结,只训练一个带有单独新关键词头的轻量级扩展分支,保留现有关键词的核心逻辑、输出和阈值。

英文摘要

Keyword spotting (KWS) models on embedded devices often need to add new keywords after deployment, but updates are difficult when original training data are unavailable and regressions on existing triggers are unacceptable. At a fixed operating point, our method reduces average new-keyword false reject rate (FRR) from 6.46 to 4.37 versus a parameter-matched separate-model baseline and outperforms parameter-efficient tuning baselines (adapters, LoRA), while using fewer multiply-accumulate operations (MACs) under the same added-parameter budget ($\leq$10k): 16.34M vs 18.45M/20.52M. We achieve this via parameter-capped modular expansion: the base network, including batch-normalization statistics and the core classifier, is frozen, and only a lightweight expansion branch with a separate new-keyword head is trained, preserving core logits, shipped outputs, and thresholds for existing keywords.

CommentsAccepted to Interspeech 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑