多通道重放语音检测的域增量学习
Domain-Incremental Learning for Multi-Channel Replay Speech Detection
浏览论文内容
中文总结 AI 辅助
针对多通道重放语音检测,提出域增量学习框架及任务特定波束形成器,在ReMASC上显著提升持续学习准确率并减少遗忘。
中文摘要 AI 辅助
重放攻击是对语音控制系统最易实施的安全威胁,而暴露这些攻击的声学线索受到攻击所处环境的强烈调制。因此,部署在现实场景中的检测器需要随时间吸收新的声学条件,理想情况下无需重新访问过去的录音,因为无限期保留语音既昂贵又受法律约束。我们将此问题构建为声学环境上的域增量学习(DIL),并提出了首个用于多通道重放语音检测的持续学习基准,在ReMASC语料库的所有24种环境排序上,使用五种随机种子评估了基于波束形成器的最先进检测器。顺序微调会严重遗忘,使先前学习环境上的错误率提高18.8个百分点。弹性权重巩固(EWC)将遗忘减半但损失了可塑性,梯度投影记忆(GPM)在统计上与朴素微调无显著差异,而所提出的任务特定波束形成器(TSB)为每个环境保留一个空间前端,显著提高了最终和增量准确率。我们进一步表明,序列中的最后一个环境主导了最终性能。代码、结果和分析可在该https URL获取。
英文摘要
Replay attacks are the most accessible threat to voice-controlled systems, and the acoustic cues that expose them are strongly modulated by the environment in which the attack is mounted. A detector deployed in the field therefore has to absorb new acoustic conditions over time, ideally without revisiting past recordings, since retaining speech indefinitely is both expensive and legally constrained. We frame this as Domain-Incremental Learning (DIL) over acoustic environments and present the first continual learning benchmark for multi-channel replay speech detection, evaluating a state-of-the-art beamformer-based detector over all 24 environment orderings of the ReMASC corpus with five seeds. Sequential fine-tuning forgets severely, raising the error rate on previously learned environments by 18.8 points. Elastic weight consolidation (EWC) halves forgetting but loses plasticity, gradient projection memory (GPM) is statistically indistinguishable from naive fine-tuning, and the proposed task-specific beamformer (TSB) that keeps one spatial front-end per environment significantly improves final and incremental accuracy. We further show that the last environment of the sequence dominates final performance. Code, results, and analysis are available at https://github.com/michaelneri/replay-speech-continual.
发表机构
- Tampere University(坦佩雷大学)
机构由 AI 辅助整理,请以论文原文为准。