发表机构
Analog Devices, Inc.; University of California, Los Angeles (UCLA)(亚德诺半导体公司; 加州大学洛杉矶分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究在严格约束下的语音活动检测问题,提出kiloVAD,利用标准Mel特征、仅CNN层等进行嵌入式推理,通过每层结构化剪枝等方法提升性能,在AVA - Speech数据集上取得新的最优结果,为因果性、可部署的VAD带来进展。
AI 中文摘要
语音活动检测(VAD)在内存、延迟和计算受限的始终在线系统中触发下游语音处理。近期紧凑模型虽准确率高,但依赖不受广泛支持的组件。本文提出kiloVAD,使用标准Mel特征、仅CNN层及可调上下文/频谱参数用于嵌入式推理。引入每层结构化剪枝、自蒸馏和基于角度的量化感知训练,性能比标准量化感知训练高1 - 4%。在因果条件下每帧评估,kiloVAD在AVA - Speech数据集上以2.1k参数和200ms上下文达到0.850的AUC,为因果性、可部署的VAD建立了新的技术水平。
英文摘要
Voice activity detection (VAD) triggers downstream speech processing in always-on systems under strict memory, latency, and compute constraints. Recent compact models report strong accuracy but rely on components that are not widely supported: learnable filterbanks, recurrent layers, or non-causal post-processing. We propose kiloVAD, designed for embedded inference using standard Mel features, CNN-only layers, and tunable context/spectral parameters. We introduce per-layer structured pruning with self-distillation and angle-based quantization-aware training (QAT) that outperforms standard QAT by 1-4%. Evaluated per-frame under causal conditions, kiloVAD achieves 0.850 AUC on AVA-Speech with 2.1 k parameters and 200 ms context, establishing a new state of the art for causal, deployment-ready VAD.
CommentsAccepted for publication at INTERSPEECH 2026