arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.24391cs.LG

NAVIR:面向边缘硬件鲁棒人机交互的神经形态音视频语音识别

NAVIR: Neuromorphic Audio-Visual Speech Recognition for Robust Human-Robot Interaction on Edge Hardware

Leonidas Delimpasis, Panagiota Moraiti, Antonis Porichis, Panos Chatzakos, Michail Karamousadakis

首次发表
浏览论文内容

中文总结 AI 辅助

NAVIR是首个在神经形态边缘硬件上运行的端到端音视频语音识别系统,通过分解时空编码和量化感知训练,在噪声下显著降低词错误率并实现高能效。

中文摘要 AI 辅助

工业环境中的语音控制交互受到声学噪声的阻碍,这会严重降低仅依赖音频的语音识别性能。音视频语音识别(AVSR)通过将唇动线索与音频流融合来解决这一问题,但最先进的流水线依赖于三维卷积、循环单元和注意力模块,这些超出了典型边缘设备的预算。我们提出了NAVIR,一个端到端的AVSR系统,针对BrainChip Akida神经形态处理器设计,该处理器原生仅支持顺序的二维卷积推理。该流水线将空间和时间编码分解为基于AkidaNet的独立模块:逐帧视觉编码器、时间视频编码器和频谱图音频编码器,通过轻量级预测头融合,并由受限束搜索解码。模型使用连接主义时间分类在噪声增强音频上训练,然后通过量化感知训练进行微调。在GRID基准上,量化后的音视频模型在未见说话人划分的噪声条件下达到14.0%的词错误率(WER),在重叠说话人划分上达到3.3%的WER,而仅音频基线分别为22.5%和11.8%;在特定任务的工业命令语料库上,它以1.5%的WER实现了98.6%的命令准确率。操作计数分析表明,在27.6%的平均发放率下,脉冲公式相比其人工神经网络对应物具有13倍的能量优势。板载测量显示,在唇读模型上,每次推理的能量比树莓派中央处理器低约5倍,比笔记本电脑图形处理器低100倍以上,同时维持每秒14.5次推理。据我们所知,这是首个在此类神经形态硬件上运行的完整多模态AVSR流水线。

英文摘要

Voice-controlled interaction in industrial settings is hampered by acoustic noise, which severely degrades audio-only speech recognition. Audio-visual speech recognition (AVSR) addresses this by fusing lip-motion cues with the audio stream, but state-of-the-art pipelines rely on three-dimensional convolutions, recurrent units, and attention modules that exceed the budget of typical edge devices. We present NAVIR, an end-to-end AVSR system targeting the BrainChip Akida neuromorphic processor, which natively supports only sequential two-dimensional convolutional inference. The pipeline factorises spatial and temporal encoding into separate AkidaNet-based modules: a per-frame visual encoder, a temporal video encoder, and a spectrogram audio encoder, fused by a lightweight predictor head and decoded by constrained beam search. Models are trained with connectionist temporal classification on noise-augmented audio and then fine-tuned with quantization-aware training. On the GRID benchmark, the quantized audio-visual model reaches 14.0% word error rate (WER) under noise on the unseen-speaker split and 3.3% WER on the overlapped-speaker split, against 22.5% and 11.8% for audio-only baselines, and it attains 98.6% command accuracy at 1.5% WER on a task-specific industrial-command corpus. Operation-count analysis indicates a 13-fold energy advantage of the spiking formulation over its artificial neural network counterpart at 27.6% mean firing rate. On-board measurements show roughly 5-fold lower energy per inference than a Raspberry Pi central processing unit on the lip-reading model, and over 100-fold lower than a laptop graphics processing unit, while sustaining 14.5 inferences per second. To the best of our knowledge, this is the first complete multimodal AVSR pipeline running on neuromorphic hardware of this class.

发表机构

  • Plaixus Ltd.(Plaixus 有限公司)
  • Tech Hive Labs(Tech Hive 实验室)
  • University of Essex(埃塞克斯大学)

机构由 AI 辅助整理,请以论文原文为准。

↑