归一化还是条件化?无人机旋翼自噪声下机载关键词识别的噪声基底前端
Normalise or condition? Noise-floor front-ends for on-board keyword spotting under UAV rotor ego-noise
浏览论文内容
中文总结 AI 辅助
本研究针对无人机旋翼自噪声下的关键词识别,比较归一化与条件化前端,发现条件化前端在保持低误警率的同时提升检测性能,且计算开销极低。
中文摘要 AI 辅助
小型多旋翼无人机机身麦克风受旋翼自噪声主导,导致语音飞行指令以负信噪比(SNR)到达。我们研究了在真实自噪声下,针对十个单词指令词汇的小足迹关键词识别(KWS),在一架四旋翼上训练,在另一架上测试。除了每片段准确率,我们还在4.4小时连续旋翼噪声上测量流式误警率。我们比较了经典噪声鲁棒前端(CMN、PCEN、谱减法)、测试时自适应,以及两种跟踪决策窗口前两秒每频带自噪声基底的前端,一种采用减法(归一化),另一种将其作为第二输入通道馈入网络(条件化)。按片段来看,所有前端表现相似:平均提升2-3个百分点,在-15 dB时最高提升14个百分点。在连续旋翼噪声上,它们差异显著。归一化前端在相同阈值下比普通log-mel触发约多十倍,并在每小时一次误警的预算下低于该阈值(0 dB时下降9个百分点)。条件化保持基线的误警率,并将其增益转化为检测(三个种子在-10 dB时提升8个百分点);在PCEN之上,它提供了最佳的每片段准确率和误警率,且一个电平锚定变体对麦克风增益不变。真实无人机+干扰源录音暴露了剩余的失败模式,即环境声音和旁观者语音,训练负样本将其减半。一个脚本可在NVIDIA Jetson Orin NX飞行计算机上复现所有设备端数字,完整流程在CPU上每100毫秒跳变耗时3毫秒。
英文摘要
A microphone on the airframe of a small multi-rotor UAV is dominated by rotor ego-noise, so spoken flight commands arrive at negative signal-to-noise ratio (SNR). We study small-footprint keyword spotting (KWS) for a ten-word command vocabulary under real ego-noise, training on one quadrotor and testing on another. Besides per-clip accuracy we measure the streaming false-alarm rate on 4.4 h of continuous rotor noise. We compare classical noise-robust front-ends (CMN, PCEN, spectral subtraction), test-time adaptation, and two front-ends that track the per-band ego-noise floor over the two seconds preceding the decision window and either subtract it (normalisation) or feed it to the network as a second input channel (conditioning). Per clip, all front-ends look alike: +2-3 points on average, up to +14 at -15 dB. On continuous rotor noise they differ sharply. Normalising front-ends fire about ten times more often than plain log-mel at the same threshold and end up below it at a budget of one false alarm per hour (-9 points at 0 dB). Conditioning keeps the baseline's false-alarm rate and turns its gain into detections (+8 points at -10 dB over three seeds); on top of PCEN it gives the best per-clip accuracy and false-alarm rate, and a level-anchored variant is also invariant to the microphone gain. Real drone+interferer recordings expose the remaining failure mode, environmental sounds and bystander speech, which training negatives halve. A single script reproduces all on-device numbers on the NVIDIA Jetson Orin NX flight computer, where the complete pipeline costs 3 ms per 100 ms hop on the CPU.