CircleMatch:基于循环时间统计的原型匹配用于微型关键词唤醒
CircleMatch: Prototype Matching with Circular Temporal Statistics for Tiny Keyword Spotting
浏览论文内容
中文总结 AI 辅助
CircleMatch通过原型匹配与循环时间统计,以约1k-7k参数的微型模型在多个关键词唤醒数据集上实现高精度,并展现出平移等变性与时间压缩适应性。
中文摘要 AI 辅助
关键词唤醒(KWS)是在语音中识别预定义词的任务,是语音设备的核心能力。在不同词汇量下,于严格的参数预算内实现高KWS精度仍然具有挑战性。我们提出了CircleMatch,一种用极少参数实现KWS的匹配框架。其编码器独立压缩频带并将其融合为帧特征。这些特征随后与学习到的类特定原型进行匹配,以生成时间响应曲线。无参数的循环聚合将时间编码为角度,并汇总响应分布和相对时序以进行分类。我们开发了四种微型变体:Circle-D4、Circle-D8、Circle-D16和Circle-D32,在12类设置中参数规模约为1k至7k。在Speech Commands v1/v2以及多语言口语词汇语料库的英语和西班牙语Micro子集上,使用多个随机种子进行的实验表明,微型模型具有竞争力的精度。我们的定性分析进一步表明原型响应具有近似平移等变性,并能适应时间压缩。代码和模型权重可在该https URL获取。
英文摘要
Keyword spotting (KWS), the task of identifying predefined words in speech, is a core capability of voice-enabled devices. Achieving high KWS accuracy under tight parameter budgets across different vocabulary sizes remains challenging. We present CircleMatch, a matching framework enabling KWS with very few parameters. Its encoder independently compresses frequency bands and fuses them into frame features. These features are then matched against learned class-specific prototypes to produce temporal response curves. Parameter-free circular aggregation encodes time as angles and summarizes response distributions and relative timing for classification. We develop four tiny variants, Circle-D4, Circle-D8, Circle-D16, and Circle-D32, ranging from approximately 1k to 7k parameters in the 12-class setting. Experiments with multiple random seeds on Speech Commands v1/v2 and the English and Spanish Micro subsets of the Multilingual Spoken Words Corpus demonstrate competitive accuracy with tiny models. Our qualitative analysis further suggests approximate shift equivariance of prototype responses and adaptation to temporal compression. Code and model weights are available at https://github.com/ora942878/CircleMatch.
发表机构
- Shanghai Normal University(上海师范大学)
机构由 AI 辅助整理,请以论文原文为准。