Lyric:用于高效口语数字识别的波域计算
Lyric: Wave-Domain Computing for Efficient Spoken-Digit Recognition
浏览论文内容
中文总结 AI 辅助
Lyric通过波域计算在数字化前提取语音特征,利用无源声学谐振器分离频率,大幅减少数据采集和计算,在保持高准确率的同时显著降低延迟和能耗。
中文摘要 AI 辅助
低功耗语音识别需要降低音频采集、特征处理以及神经推理的成本。我们提出了Lyric,一种在数字化之前利用波域计算提取特征的语音识别前端。我们提出的系统使用无源声学谐振器按频率分离语音,并将其缓慢变化的包络提供给紧凑的时间神经网络。这将频谱滤波移入声学结构,减少了识别所需的数据和处理量。我们构建了一个原型,并在九个分类器家族上对说话者独立的AudioMNIST分割进行了口语数字识别评估。该前端获取的标量样本比16-kHz波形少32倍。在Raspberry Pi 4上的EdgeSpeechNet-A比较中,我们将预处理加推理的延迟和每次推理的估计处理器能量减少了98.6%,准确率下降了3.58个百分点,降至95.70%。
英文摘要
Low-power speech recognition requires reducing audio acquisition and feature-processing costs as well as neural inference. We present Lyric, a speech-recognition front end that uses wave-domain computing to extract features before digitization. Our proposed system uses passive acoustic resonators to separate speech by frequency and supplies their slowly varying envelopes to a compact temporal neural network. This moves spectral filtering into the acoustic structure, reducing the data and processing needed for recognition. We build a prototype and evaluate spoken-digit recognition on a speakerdisjoint AudioMNIST split across nine classifier families. The front end acquires 32 times fewer scalar samples than a 16-kHz waveform. In the EdgeSpeechNet-A comparison on a Raspberry Pi 4, we reduce preprocessing-plus-inference latency and estimated processor energy per inference by 98.6%, with a 3.58-percentage-point decrease in accuracy to 95.70%.