arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

子模型短期记忆卷积用于设备端关键词识别系统

Sub-Model Short-Term Memory Convolutions for Keyword Spotting Systems on Device

Paweł Warlewski, Artur Czeczko, Artur Szumaczuk, Grzegorz Stefański, Szymon Klimaszewski

arXiv 2609.35005首次发表:更新:

发表机构

Samsung R&D Institute Poland; Samsung AI Center Warsaw(三星研发中心波兰分部; 三星AI中心华沙)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出将STMC框架应用于模块化CNN,实现在线类LSTM推理,降低功耗和冗余计算,在关键词识别任务上分别减少82%和46%的MCPS,并达到93.8%和97.1%的准确率。

AI 中文摘要

关键词识别(KWS)随着语音控制设备的日益普及而变得越来越重要。虽然与智能手机和智能电视的语音交互已经普遍,但在可穿戴设备等资源严重受限的边缘设备上部署KWS仍然具有挑战性。这些系统必须在计算能力、内存占用和实时延迟的严格约束下满足高精度要求。在这项工作中,我们提出了STMC(短期记忆卷积)框架的一种应用,以将模块化CNN模型适配为在线、类似LSTM的推理。我们的方法降低了功耗和冗余计算,同时保持了训练CNN的稳定性和简单性。与同等频率的标准CNN执行和原始STMC相比,我们分别实现了高达82%和46%的MCPS降低。最佳配置在11类Google Speech Commands任务上达到93.8%的准确率,在相同任务的零填充数据上达到97.1%的准确率。

英文摘要

Keyword Spotting (KWS) is becoming increasingly important as voice-controlled devices grow more widespread. While voice interaction with smartphones and smart TVs is already common, deploying KWS on heavily resource-constrained edge devices such as wearables remains challenging. These systems must meet high accuracy requirements while operating under strict constraints on computational power, memory footprint, and real-time latency. In this work, we present an application of the STMC (Short-Term Memory Convolutions) framework to adapt a modular CNN model for online, LSTM-like inference. Our approach reduces power consumption and redundant computations while maintaining the stability and simplicity of training CNNs. We achieve up to 82% and 46% MCPS reduction compared to equivalently frequent standard CNN execution and vanilla STMC, respectively. The best configuration achieves 93.8% accuracy on the 11-class Google Speech Commands task and 97.1% on the same task with zero-padded data.

CommentsInterspeech 2026, 5 pages, 2 figures

Journal refProc. Interspeech 2026, 4077-4081

DOI:10.21437/Interspeech.2026-1343

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑